Overview:
The Data Engineer builds and maintains the source adapters and normalization logic that translate raw data from disparate systems into a common risk-signal schema. This role focuses on reliable ingestion and transformation—turning heterogeneous legacy inputs (APIs, feeds, databases, files, and event streams) into consistent, high-quality signals that downstream scoring and adjudication workflows can trust.
This position is remote but will require travel in the DMV area.
What you will do:
- Build source adapters/connectors to ingest data from APIs, legacy systems, databases, and event streams
- Develop normalization and mapping logic to translate source-specific fields into the common risk-signal schema (including validation, enrichment, and standardization)
- Implement ETL/ELT pipelines with strong engineering rigor: testing, observability, error handling, retries, and backfills
- Produce and consume streaming events (e.g., Kafka topics) to support near-real-time signal delivery and downstream processing
- Partner with data architecture and domain SMEs to define and maintain data contracts, mappings, and lineage from source to normalized signal
- Ensure data quality and consistency (deduplication patterns, schema evolution handling, and reconciliation against source systems)
- Optimize pipeline performance and reliability (throughput, latency, and scalable processing patterns)
- Create and maintain technical documentation for adapters, transformations, and operational runbooks
- Some travel may be required within the DMV area
What you need to have:
- Clearance: Must maintain an active Top Secret security clearance
- Bachelor's Degree and 8 to 10 years of experience; Master's Degree and 6 to 8 years of experience
- 3–5 years of experience in data engineering, including building production-grade ingestion and transformation pipelines.
- Strong experience with API integrations and ETL/ELT development in complex environments.
- Experience integrating heterogeneous and/or legacy systems with inconsistent schemas and data quality.
- Proficiency in Python or Java for building data services and transformation logic.
- Solid SQL skills and working familiarity with NoSQL data stores.
- Experience with REST/API frameworks and building maintainable, well-tested integration services.
- Hands-on experience producing/consuming events in Kafka (producers/consumers) or an equivalent event streaming platform
- IC/DoD experience
What we'd like you to have:
Tools & Technical Environment (Preferred/Used)
- Python or Java
- REST/API frameworks
- Kafka producers/consumers
- SQL and NoSQL databases
Key Behavioral Competencies
- Engineering discipline: writes maintainable, testable code and builds robust pipelines that handle edge cases.
- Curiosity and persistence: digs into messy source data and drives it to consistent outcomes.
- Collaboration: works effectively across data architecture, scoring/analytics, and application teams.
- Operational mindset: builds pipelines that are observable, debuggable, and supportable in production.
About BigBear.ai:
BigBear.ai is a leading provider of AI-powered decision intelligence solutions for national security, supply chain management, and digital identity. Customers and partners rely on Bigbear.ai’s predictive analytics capabilities in highly complex, distributed, mission-based operating environments. Headquartered in McLean, Virginia, BigBear.ai is a public company traded on the NYSE under the symbol BBAI. For more information, visit https://bigbear.ai/ and follow BigBear.ai on LinkedIn: @BigBear.ai and X: @BigBearai.
BigBear.ai is an Equal opportunity employer all protected groups, including protected veterans and individuals with disabilities.