| Location | Alpharetta, GA |
Java Data Engineer
[City, State] / Remote / Hybrid
Full-time
We are seeking a highly motivated Java Data Engineer to design, develop, and maintain scalable data pipelines and distributed data processing systems. The ideal candidate will have expertise in Java, big data technologies, cloud platforms, and ETL development. You will work closely with data scientists, software engineers, and business stakeholders to build robust data solutions that enable analytics and business intelligence.
Design, develop, and maintain scalable data pipelines using Java.
Build and optimize ETL/ELT workflows for processing large datasets.
Develop high-performance data ingestion, transformation, and integration solutions.
Design and maintain distributed data processing applications using Spark or Hadoop.
Develop RESTful APIs for data services and integrations.
Optimize SQL queries and database performance.
Implement data quality, validation, and monitoring processes.
Collaborate with cross-functional teams to understand business data requirements.
Troubleshoot production data issues and optimize pipeline performance.
Ensure data security, governance, and compliance standards are followed.
Participate in Agile development, code reviews, and architecture discussions.
Maintain technical documentation and best practices.
Bachelor's degree in Computer Science, Information Technology, Data Engineering, or a related field.
3+ years of experience in Java development.
Experience with data engineering or ETL development.
Strong proficiency in Java (Java 8/11/17).
Experience with Spring Boot.
Strong SQL and database design skills.
Experience working with relational and NoSQL databases.
Experience building REST APIs.
Familiarity with Linux/Unix environments.
Experience with Git and version control.
Knowledge of Agile/Scrum methodologies.
Experience with Apache Spark, Hadoop, or Apache Flink.
Knowledge of Kafka or other messaging platforms.
Experience with cloud platforms (AWS, Azure, or GCP).
Experience with data lakes and data warehouses.
Familiarity with containerization using Docker and Kubernetes.
Experience with Airflow or workflow orchestration tools.
Knowledge of data modeling and dimensional modeling.
Experience with CI/CD pipelines.
Java 8/11/17
SQL
Python (Preferred)
Scala (Nice to Have)
Spring Boot
Spring Batch
Spring Data
Hibernate
Apache Spark
Hadoop
Hive
Apache Flink
Apache Kafka
Apache NiFi
Apache Airflow
Talend
Informatica
Spring Batch
PostgreSQL
MySQL
Oracle
SQL Server
MongoDB
Cassandra
Redis
AWS
S3
EMR
Glue
Redshift
Lambda
Microsoft Azure
Azure Data Factory
Azure Synapse
Azure Blob Storage
Google Cloud Platform
BigQuery
Dataflow
Cloud Storage
Snowflake
Amazon Redshift
Google BigQuery
Azure Synapse Analytics
Docker
Kubernetes
Jenkins
GitHub Actions
Azure DevOps
Git
GitHub
GitLab
Bitbucket
Splunk
ELK Stack (Elasticsearch, Logstash, Kibana)
Grafana
Prometheus
Strong analytical and problem-solving abilities.
Excellent communication and collaboration skills.
Ability to work independently and in a team environment.
Strong attention to detail.
Good organizational and time management skills.
Ability to manage multiple projects and deadlines.
Experience building enterprise-scale data platforms.
Experience processing structured and unstructured data.
Knowledge of streaming data architectures.
Experience optimizing large-scale data pipelines.
Familiarity with data governance and security practices.
Experience working in Agile development teams.
Apache Beam
Delta Lake
Databricks
Iceberg
Hudi
Terraform
dbt
OAuth/JWT Authentication
Microservices Architecture
Event-Driven Architecture
Machine Learning data pipelines
Real-time analytics platforms
Experience: 3 8+ Years
Notice Period: Immediate to 30 Days Preferred