Design, develop, and maintain scalable batch and streaming data pipelines supporting customer migration activities. Create reusable frameworks and components to improve data pipeline performance, reliability, and maintainability.
Numbers & Facts
Location
Wilmington, DE
Description
Key Responsibilities
Design, develop, and maintain scalable batch and streaming data pipelines supporting customer migration activities.
Build reliable data ingestion, transformation, validation, and delivery processes across distributed systems.
Develop production-quality applications and services using Java and Spring Boot.
Use Apache Spark and PySpark to process, transform, and analyze large datasets.
Develop and maintain AWS-based data solutions using services such as AWS Glue, AWS Lambda, and Amazon S3.
Create reusable frameworks and components to improve data pipeline performance, reliability, and maintainability.
Implement data quality checks, validation rules, reconciliation processes, and error-handling mechanisms.
Support the migration of customer and account data while ensuring accuracy, completeness, security, and regulatory compliance.
Integrate data pipelines with internal applications, APIs, databases, messaging systems, and downstream platforms.
Develop RESTful APIs and supporting services where required.
Build automated unit, integration, and functional tests for data pipelines and Java applications.
Participate in code reviews, design discussions, technical documentation, and development standards initiatives.
Implement and maintain CI/CD pipelines to automate build, test, deployment, and release processes.
Monitor pipeline execution, troubleshoot failures, and perform root-cause analysis for production issues.
Optimize Spark jobs, data-processing workflows, and cloud resources for improved performance and cost efficiency.
Collaborate with product managers, architects, application developers, QA engineers, DevOps teams, and business stakeholders.
Support deployment activities across development, test, staging, and production environments.
Contribute to technical design documents, operational runbooks, support procedures, and knowledge-sharing sessions.
Follow Capital One security, risk-management, data-governance, and software-development standards.
Required Qualifications and Technical Skills
Strong professional experience in data engineering and data pipeline development.
Hands-on Java development experience in enterprise applications.
Experience developing applications and services using Spring Boot.
Strong experience with Apache Spark and PySpark.
Hands-on experience with AWS Glue for data integration and transformation.
Experience developing or supporting AWS Lambda functions.
Strong experience working with Amazon S3 for data storage and data-lake solutions.
Experience designing, developing, testing, and supporting scalable data pipelines.
Experience with CI/CD pipelines and automated software delivery practices.
Strong understanding of data ingestion, transformation, validation, reconciliation, and error handling.
Experience troubleshooting production data issues and supporting mission-critical applications.
Ability to work in a fast-paced, highly collaborative enterprise technology environment.
Preferred Qualifications
Professional experience developing data solutions using Python.
Experience developing RESTful APIs and microservices.
Exposure to artificial intelligence, machine learning, or AI-enabled data platforms.
Experience with Kafka or other messaging and event-streaming technologies.
Experience supporting customer, banking, payments, credit-card, or financial-services data.
Previous Capital One experience.
Experience with cloud-native application development and distributed systems.
Familiarity with data governance, security, privacy, and regulatory requirements in financial services.
Experience with infrastructure automation, containerized applications, or Kubernetes is a plus.