| Location | Atlanta, GA |
Back to Search
Senior Data Software Engineer, Spark, Java, Scala
Remote in Georgia, & 4 others
Data Software Engineering
apply
FacebookLinkedInSend via email
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a job
Location-specific conditions & benefits*
Choose an option
We are seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This is not a data-science or analytics position centered on ad-hoc data exploration; instead, the role focuses on building software, data processing jobs, and data pipelines consumed by internal and external partners. Data is our main product and first-class citizen, and we value correct, high-quality data as much as clean and maintainable code.
Responsibilities
Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development
Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages
Fix bugs in code and correct data caused by incorrect logic or implementation
Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions
Monitor and troubleshoot production issues with pipelines owned by the team
Develop and adopt data quality checks to monitor data issues in the systems
Scope and plan new development, including assessing level of effort and providing timelines
Maintain tickets hygiene in Radar (ticketing system)
Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation
Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines
Requirements
3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark
Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL
Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3
Excellent knowledge of Scala or Java
Understanding of batch processing and ETL principles in Data Warehouses
Familiarity with data completeness signals and orchestration
Knowledge of approaches for historical reprocessing and data correction
Skills in handling bad data and late data in inputs and outputs
Understanding of schema migrations and datasets evolution
Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers
Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more
Nice to have
Understanding of functional programming ideas and principles
Experience in building and using web services
Familiarity with any of Teradata, Vertica, Oracle, Tableau
Skills in Spark Streaming and Kafka
Knowledge of Apache Iceberg, Trino (Presto), Druid, Cassandra, or Blob storage like AWS
Experience with Splunk
Experience with Snowflake