| Location | Chicago, IL |
Job Description
Job Title: Sr. Data Engineer with AI Experience
Experience: - 8+ Years
Location: -Chicago, IL (Hybrid 3 days' Work from Office)
In person interview required
Skills: Azure DataBricks ADF, DataBricks, Python, Pyspark, ETL, SQL.
Roles and Responsibilities: -
Rebuild Epic Caboodle extraction from a transformation-heavy pattern to clean incremental raw ingestion - no joins, no temp tables, no business logic at the source
Implement incremental / change-data-capture ingestion into the landing and raw layers on Azure and Databricks
Move member and facility mapping downstream, out of the acquisition step. Analyse and optimize existing SQL.
Work through the existing per-member script inventory with the Epic SME to map current Caboodle table access patterns - Identify redundant table access.
Establish the baseline table-touch count and measure the reduction the new design delivers
Implement the consolidation of transformations from three stages into two, moving redundant joins and source-specific logic into the layout-converter stage alongside the final data-model conversion
Restructure processing to run by source system rather than per member - Keep filtering out of the intermediate layer so the domain-refined mart remains the source of truth
Validate agent-generated output. Review, execute, and validate the SQL and metadata the agents generate - you are the quality gate between agent proposal and production execution
Build and run reconciliation harnesses comparing new-path output against current production extracts - Confirm downstream consumers (Databricks, SQL Server, Vertica) are not adversely affected by schema or semantic changes
Build tooling the agents depend on Data profiling routines the agents call as tools - null rates, cardinality, distributions, date coverage, code-system detection - Metadata and schema extraction jobs that seed the project's knowledge graph - Instrumentation for cost, run time, and table-touch telemetry.
Educational Qualifications: -
Engineering Degree BE/ME/BTech/MTech/BSc/MSc.
Technical certification in multiple technologies is desirable.
Skills: -
Mandatory skills
5+ years in data engineering, with recent hands-on delivery ownership.
Strong Databricks and PySpark jobs, workflows, Delta Lake, performance tuning
Expert SQL including the ability to read unfamiliar, poorly documented SQL at volume and reason about what it does and what it costs
Demonstrable query cost and performance optimization experience, this is a core requirement on this engagement, not a bonus
Incremental / CDC ingestion pattern** design and implementation - Azure data services - ADLS / Blob Storage, and Databricks on Azure.
Unity Catalog or comparable data governance and cataloguing experience - Data profiling and reconciliation testing proving two pipelines produce equivalent output
Comfort working with PHI-scoped healthcare data and the access controls that implies
Ability to work independently against ambiguous inputs and drive clarification, in a small team on a hard deadline
Preferred - Epic EHR data experience, Caboodle or Clarity schema familiarity is a significant advantage
Healthcare data domain knowledge: clinical coding systems (ICD, CPT, SNOMED, LOINC), encounter and procedure data models
Workflow orchestration experience - Orkes / Netflix Conductor, or transferable experience with Airflow, Dagster, or similar - SQL Server and/or Vertica exposure, for downstream compatibility validation
Experience working alongside AI/LLM-generated code or configuration - Knowledge-graph or metadata-management exposure (Stardog, RDF/SPARQL, or similar)
Experience with AI-assisted development tooling and spec-driven delivery practices
Someone with EPIC experience / knowledge is good to have