Senior AWS + Databricks Analyst (Data Architect Level)

DCM INFOTECH LIMITED

  • Albany, NY
  • Today

    Highlights

    Technical Skills Matrix Domain Required Technologies Cloud Platform Lakehouse AWS — S3, Glue, Lambda, Step Functions, EMR, Athena, Redshift, DMS, IAM, KMS, CloudWatch, VPC Databricks, Delta Lake, Unity Catalog, Delta Live Tables, Databricks SQL, Photon Processing Data Modeling Apache Spark, PySpark, Spark SQL, Python Streaming Governance Dimensional modeling, medallion architecture, data vault (nice to have) Kinesis, Kafka/MSK, Auto Loader, structured streaming Unity Catalog, Lake Formation, RBAC, data lineage, PII masking/tokenization DevOps Reporting Git, Terraform/CloudFormation, CI/CD pipelines, automated testing for data Power BI / Tableau integration, Databricks SQL warehouse Position Summary The Senior AWS + Databricks Analyst will serve as the senior technical authority for the agency's cloud data platform, owning the end-to-end architecture of large scale analytics workloads on AWS with Databricks as the primary lakehouse platform.

    Numbers & Facts

    LocationAlbany, NY

    Description

    Position Summary The Senior AWS + Databricks Analyst will serve as the senior technical authority for the agency's cloud data platform, owning the end-to-end architecture of large scale analytics workloads on AWS with Databricks as the primary lakehouse platform. This is not an execution-only role — the resource is expected to define target-state architecture, set data engineering standards, lead design reviews, and advise agency leadership on platform direction, cost governance, and modernization roadmap. The candidate will work directly with agency program staff, data stewards, and existing vendor teams to migrate and modernize legacy data assets, build governed data products, and operationalize analytics for statewide reporting and program decision-making. Key Responsibilities Own the target-state architecture for the agency's AWS-based data platform, including landing, curated, and consumption zones on S3 with Databricks Delta Lake as the storage and processing standard. Design and document medallion (bronze/silver/gold) architectures, data models, and ingestion patterns for structured, semi structured, and streaming sources. Architect and review Databricks workloads: Delta Live Tables, Unity Catalog governance, workflows/jobs orchestration, cluster policies, Photon optimization, and performance tuning. Define and enforce data governance, lineage, access control, and PII handling standards using Unity Catalog, AWS Lake Formation, IAM, and KMS. Lead the assessment and migration of legacy on premises data platforms (SQL Server, Oracle, DB2, mainframe extracts, SSIS/Informatica) to the AWS/Databricks lakehouse. Build and review production-grade pipelines using PySpark, Spark SQL, and Python; establish reusable frameworks, coding standards, and CI/CD practices. Architect the surrounding AWS ecosystem: S3, Glue, Lambda, Step Functions, EMR, Kinesis/MSK, Redshift, Athena, DMS, CloudWatch, and VPC/networking for secure data movement. Establish FinOps discipline — cluster right-sizing, job scheduling, autoscaling policies, spot strategy, and chargeback reporting to control platform spend. Conduct architecture and code reviews; mentor agency staff and vendor engineers; produce design documents, runbooks, and knowledge transfer artifacts suitable for state audit. Support analytics and BI consumers (Power BI, Tableau, or agency standard) with semantic layer design and query performance optimization. Participate in agency change control, security review, and NYS ITS architecture review board processes as required. Mandatory Qualifications Candidates must demonstrate the following on the resume with dates, client names, and role context. Experience is stated in months to align with NYS submission requirements. 1. 120 months of overall professional experience in data engineering, data warehousing, or data platform architecture. 2. 72 months of hands-on experience architecting and delivering data solutions on Amazon Web Services, including S3, Glue, Lambda, IAM, and at least one AWS analytics service (Redshift, EMR, Athena, or Kinesis). 3. 48 months of hands-on experience with Databricks, including Delta Lake, notebooks, jobs/workflows, and cluster administration. 4. 60 months of hands-on development experience with Apache Spark (PySpark and/or Spark SQL) at production scale. 5. 72 months of experience with SQL and dimensional/relational data modeling (star schema, slowly changing dimensions, normalization). 6. 36 months of experience in a lead or architect capacity, including responsibility for solution design, design documentation, and technical review of other engineers' work. 7. 24 months of experience implementing data governance and security controls — role-based access, encryption at rest/in transit, PII masking, or catalog-based governance (Unity Catalog, Lake Formation, Collibra, or equivalent). 8. 24 months of experience with CI/CD and infrastructure-as-code for data platforms (Git, Terraform or CloudFormation, Azure DevOps/Jenkins/GitHub Actions). 9. Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field. (Additional four years of directly relevant experience may substitute for the degree.) 10. Demonstrated experience producing formal architecture deliverables — target-state diagrams, data flow documentation, and design decision records. Desirable Qualifications Prior experience delivering for a state, local, or federal government agency, ideally New York State (ITS, OGS, DOH, OTDA, DOL, Tax & Finance, or similar). Databricks Certified Data Engineer Professional or Databricks Certified Data Analyst Associate. AWS Certified Solutions Architect – Professional or AWS Certified Data Analytics – Specialty. Experience with Unity Catalog rollout and migration from Hive metastore. Experience with Delta Live Tables and streaming ingestion (Kinesis, Kafka/MSK, Auto Loader). Experience with MLflow or supporting data science/ML workloads on Databricks. Experience migrating from Informatica, DataStage, SSIS, or Ab Initio to Spark-native pipelines. Familiarity with NYS security standards (NYS-S13 001 et seq.), NIST 800-53, or FedRAMP-aligned control environments. Experience with public-sector data domains: health and human services, tax, workforce, transportation, or benefits eligibility. Working knowledge of BI tooling — Power BI, Tableau, or Cognos — and semantic layer design. Technical Skills Matrix Domain Required Technologies Cloud Platform Lakehouse AWS — S3, Glue, Lambda, Step Functions, EMR, Athena, Redshift, DMS, IAM, KMS, CloudWatch, VPC Databricks, Delta Lake, Unity Catalog, Delta Live Tables, Databricks SQL, Photon Processing Data Modeling Apache Spark, PySpark, Spark SQL, Python Streaming Governance Dimensional modeling, medallion architecture, data vault (nice to have) Kinesis, Kafka/MSK, Auto Loader, structured streaming Unity Catalog, Lake Formation, RBAC, data lineage, PII masking/tokenization DevOps Reporting Git, Terraform/CloudFormation, CI/CD pipelines, automated testing for data Power BI / Tableau integration, Databricks SQL warehouse

    Similar Jobs

    See more jobs