Senior Data Engineer

Spark Tek Inc

San Jose, CA

JOB DETAILS
SKILLS
Access Control, Amazon Simple Storage Service (S3), Amazon Web Services (AWS), Apache, Apache Hadoop, Apache Kafka, Apache Spark, Application Programming Interface (API), Architectural Services, Best Practices, Centers for Disease Control and Prevention (CDC), Cloud Computing, Code Reviews, Communication Skills, Computer Programming, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Cross-Functional, Data Analysis, Data Management, Data Modeling, Data Processing, Data Quality, Data Science, Data Sets, Data Warehousing, Database Extract Transform and Load (ETL), DevOps, Dimensional Modeling, Docker, GCP (Good Clinical Practices), Git, Information/Data Security (InfoSec), Java, Machine Learning, Mentoring, Microsoft Windows Azure, Privacy Controls, Python Programming/Scripting Language, Quality Monitoring, SQL (Structured Query Language), Scala Programming Language, Scalable System Development, Snowflake Schema, Software Engineering, Source Code/Configuration Management (SCM), Star Schema, Startup, Streaming Technology, Technical Leadership, Technical/Engineering Design, Transformation Tools, Use Cases
LOCATION
San Jose, CA
POSTED
1 day ago

Job Title: Senior Data Engineer

Location: San Jose, CA/Dallas, TX - Hybrid

Job Type: Contract

Job Description:

  • We are looking for a Senior Data Engineer with 7–8 years of experience to design, build, and maintain scalable data infrastructure and pipelines. You will work closely with data scientists, analysts, and software engineering teams to ensure data is reliable, accessible, and optimized for analytics and decision-making. This is a senior individual-contributor role with strong ownership over architecture decisions and mentorship of junior engineers.

Key Responsibilities:

  • Design, build, and maintain robust, scalable ETL/ELT pipelines to ingest data from diverse sources (databases, APIs, streaming platforms, third-party systems).
  • Architect and optimize data warehouse/lakehouse solutions (e.g., Snowflake, BigQuery, Redshift, Databricks) for performance, scalability, and cost efficiency.
  • Build and maintain batch and real-time streaming data pipelines using tools such as Apache Kafka, Spark, Flink, or similar.
  • Own the design of data models (dimensional modeling, star/snowflake schemas) to support analytics, reporting, and ML use cases.
  • Implement data quality checks, monitoring, alerting, and observability across pipelines to ensure accuracy and reliability.
  • Collaborate with data scientists and analysts to understand data requirements and deliver clean, well-documented datasets.
  • Drive best practices around data governance, security, access control, and compliance (e.g., GDPR, SOC 2).
  • Optimize infrastructure costs and pipeline performance, identifying and resolving bottlenecks.
  • Mentor junior and mid-level data engineers; participate in code reviews and technical design discussions.
  • Partner with DevOps/Platform teams to manage CI/CD pipelines, infrastructure as code, and containerized deployments for data workloads.
  • Evaluate and recommend new tools, frameworks, and architectural patterns to improve the data platform.

Required Skills & Qualifications:

  • 7–8 years of hands-on experience in data engineering, backend engineering, or a related field.
  • Strong programming skills in Python and/or Scala/Java; advanced SQL proficiency required.
  • Deep experience with distributed data processing frameworks (Apache Spark, Hadoop, or similar).
  • Hands-on experience with cloud platforms (AWS, GCP, or Azure) and their data services (S3, Glue, Redshift, BigQuery, Dataflow, Data Factory, etc.).
  • Experience with workflow orchestration tools such as Apache Airflow, Dagster, or Prefect.
  • Solid understanding of data modeling concepts (dimensional modeling, normalization, CDC, SCD types).
  • Experience with streaming technologies (Kafka, Kinesis, Pub/Sub, or similar).
  • Proficiency with modern data warehouse/lakehouse platforms (Snowflake, Databricks, BigQuery, Redshift).
  • Strong understanding of data infrastructure best practices: version control (Git), CI/CD, containerization (Docker/Kubernetes), and infrastructure as code (Terraform).
  • Experience implementing data quality frameworks and monitoring/observability tools (e.g., Great Expectations, Monte Carlo, Datadog).
  • Strong understanding of data security, privacy, and compliance practices.
  • Excellent communication skills and ability to work cross-functionally with technical and non-technical stakeholders.
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field (or equivalent practical experience).

Preferred / Nice-to-Have:

  • Experience with real-time analytics and event-driven architectures.
  • Exposure to machine learning pipelines and MLOps practices.
  • Experience with dbt (data build tool) for transformation workflows.
  • Prior experience mentoring teams or leading technical projects.
  • Relevant certifications (AWS/GCP/Azure Data Engineer certifications).
  • Experience in a high-growth startup or high-scale enterprise environment.

About the Company

S

Spark Tek Inc