Lead Data Engineer Financial Data & Analytics

Saicon Consultants Inc

  • Parsippany, NJ
  • 1 day ago

    Highlights

    This role is approximately 70% data engineering and 30% analytics and is well suited for someone who enjoys both sides of the data lifecycle: building reliable production pipelines and digging into complex datasets to understand what the data is actually saying. The ideal candidate brings deep experience with Databricks, PySpark, Python, SQL, and large-scale data processing, along with exposure to financial services, payroll, economic, or other transaction-intensive data environments.

    Numbers & Facts

    LocationParsippany, NJ

    Description

    Position Overview

    We are seeking an experienced Lead Data Engineer to help build and evolve modern data platforms supporting large-scale financial, economic, and analytical workloads.

    This role is approximately 70% data engineering and 30% analytics and is well suited for someone who enjoys both sides of the data lifecycle: building reliable production pipelines and digging into complex datasets to understand what the data is actually saying.

    The ideal candidate brings deep experience with Databricks, PySpark, Python, SQL, and large-scale data processing, along with exposure to financial services, payroll, economic, or other transaction-intensive data environments. You will work closely with data scientists, analysts, business stakeholders, and engineering teams to develop scalable data products that support research, reporting, advanced analytics, and machine learning initiatives.

    This is a hands-on technical role requiring someone who can independently solve complex data problems, challenge questionable results, and translate analytical concepts into reliable production solutions.

    Key Responsibilities

    Data Engineering & Platform Development

    • Design, develop, and maintain scalable data pipelines for ingesting, transforming, integrating, and distributing large and complex datasets.

    • Build and enhance cloud-based data platforms using Databricks, Apache Spark, PySpark, and Delta Lake.

    • Develop ETL/ELT frameworks supporting structured, semi-structured, and unstructured data from multiple internal and external sources.

    • Create data models and curated datasets optimized for analytics, reporting, research, and machine learning use cases.

    • Engineer solutions capable of processing multi-terabyte datasets and billions of records, including high-volume time-series data.

    • Optimize Spark workloads and data structures for performance, reliability, and cost efficiency.

    • Establish and maintain strong standards around data quality, lineage, observability, metadata, and governance.

    • Troubleshoot production data issues and identify root causes across complex pipelines and datasets.

    Analytics & Data Science Partnership

    • Perform exploratory data analysis to understand new datasets, including distributions, anomalies, missing values, outliers, trends, and data drift.

    • Partner with data scientists and quantitative stakeholders to convert analytical concepts into efficient production-grade implementations.

    • Develop complex aggregations, time-window calculations, cross-sectional metrics, and feature-engineering pipelines at scale.

    • Create notebooks, dashboards, and validation processes to evaluate pipeline outputs and analytical results.

    • Independently sanity-check results and investigate situations where data does not align with expected business or statistical patterns.

    • Use pandas and NumPy for exploratory and ad-hoc analysis while leveraging PySpark for large-scale production workloads.

    • Support research involving financial, economic, transactional, or other large-scale time-series datasets.

    Data Architecture & Governance

    • Contribute to the design and evolution of a modern lakehouse architecture.

    • Participate in technical architecture discussions involving data organization, scalability, performance, security, and governance.

    • Apply concepts such as medallion architecture, data products, and data mesh where appropriate.

    • Work with Unity Catalog or similar technologies to manage data access, governance, catalog structures, and schemas.

    • Implement appropriate security and access controls for sensitive business and financial information.

    • Help establish reusable engineering standards and patterns across the broader data environment.

    DevOps & Automation

    • Build automated deployment processes for data pipelines, notebooks, jobs, and platform components.

    • Implement CI/CD practices using technologies such as Jenkins, Bitbucket Pipelines, Git, and Databricks Asset Bundles or comparable tooling.

    • Support automated testing, deployment validation, monitoring, and rollback processes.

    • Partner with infrastructure and platform teams on environment configuration and infrastructure-as-code initiatives.

    AI-Assisted Engineering

    • Use modern AI development tools to accelerate coding, troubleshooting, documentation, testing, and analytical workflows.

    • Critically evaluate AI-generated solutions rather than relying on generated code without validation.

    • Review AI-generated code for accuracy, scalability, performance, security, and maintainability.

    • Identify opportunities to incorporate AI-assisted development practices into data engineering workflows.

    Required Qualifications

    • Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, Economics, Finance, or a related discipline.

    • 5+ years of professional experience in data engineering, data platform engineering, or a closely related field.

    • Strong hands-on experience with:

      • Databricks

      • Apache Spark / PySpark

      • Python

      • SQL

      • Delta Lake

      • Data modeling

      • ETL/ELT development

    • Experience engineering and optimizing very large datasets, including multi-terabyte or billion-record environments.

    • Strong knowledge of Databricks architecture and platform capabilities, including areas such as:

      • Unity Catalog

      • Delta Lake optimization and versioning

      • Databricks Workflows

      • Deployment and packaging frameworks

    • Experience implementing CI/CD processes for data engineering workloads.

    • Working knowledge of pandas and NumPy for data exploration and analysis.

    • Ability to perform exploratory data analysis and understand basic statistical concepts such as distributions, correlation, outliers, trends, and time-series behavior.

    • Experience implementing data quality, reconciliation, validation, and monitoring processes.

    • Ability to participate meaningfully in data architecture and platform design discussions.

    • Experience working with financial, payroll, banking, investment, insurance, or similarly complex transactional datasets.

    • Strong analytical skills with the ability to question unexpected results and investigate underlying data issues.

    • Strong communication and documentation skills.

    Preferred Qualifications

    • Experience working with macroeconomic, financial markets, payroll, banking, investment, or alternative datasets.

    • Experience designing or working within modern lakehouse architectures.

    • Knowledge of medallion architecture and data mesh concepts.

    • Experience with Kafka, Spark Structured Streaming, or other event-driven data architectures.

    • Experience processing geographic or census-related datasets.

    • Familiarity with infrastructure-as-code technologies such as Terraform.

    • Databricks certifications at the Associate or Professional level.

    • Experience modernizing or migrating legacy data platforms to Databricks or similar cloud-native architectures.

    • Experience supporting machine learning, predictive analytics, or AI initiatives.

    • Experience with visualization and reporting technologies such as Power BI, Tableau, or Databricks dashboards.

    • Scala experience is a plus.

    Technical Environment

    Programming & Data Processing:
    Python, SQL, PySpark, Apache Spark, pandas, NumPy; Scala is a plus

    Data Platform:
    Databricks, Delta Lake, Unity Catalog, lakehouse architecture

    Data & Messaging Technologies:
    Relational databases, Delta tables, NoSQL platforms, Kafka

    DevOps & Automation:
    Git, CI/CD pipelines, Jenkins, Bitbucket or similar source-control/deployment platforms, Terraform

    Analytics & Visualization:
    Power BI, Databricks dashboards, Python-based visualization tools

    What We're Looking For

    The strongest candidates will combine the mindset of a data engineer with the curiosity of an analyst. You should be comfortable taking ownership of a production pipeline one day and investigating an unfamiliar dataset the next.

    You should be someone who:

    • Builds scalable solutions without losing sight of what the underlying data represents.

    • Can determine whether an analytical result makes sense rather than simply trusting successful code execution.

    • Understands how to balance performance, maintainability, data quality, and delivery speed.

    • Is comfortable partnering with data scientists, analysts, engineers, and business stakeholders.

    • Can independently explore a dataset, identify meaningful patterns or problems, and clearly communicate findings.

    • Uses AI development tools as an accelerator while maintaining ownership of the accuracy and quality of the final solution.

    • Approaches complex data problems with curiosity, sound judgment, and strong attention to detail.

    Similar Jobs

    Neotecra

    Data Engineer

    • New York, NY
    13 days ago
    See more jobs