Tekfortune is a fast-growing consulting firm specialized in permanent, contract & project-based staffing services for world s leading organizations in a broad range of industries. In this quickly changing economic landscape, virtual recruiting and remote work are critical for the future of work. To support the active project demands and skills gaps, our staffing experts can help you find the best job for you.
Data Engineer
Remote working EST
Contract to hire Overview: Have multiple concurrent initiatives, all of which requiring Data Engineering support to work closely with their Data Science team: compute infrastructure migration, tech debt reduction, expansion of tool into new use case.
2 products they develop and maintain on his team, both in support of Caremark and PBM as a business and both in support of underwriting and client consulting stakeholder:
- ACCT Price Benchmarking Tool - Suite of tools to help them understand Caremark deal values. Deals consist of ~50 parameters to define key pricing terms, driving complexity to determine whether a deal is "good" or "bad." ACCT takes pricing components and other parameters, data from clients similar to one another, and generates a comparison.
- Excel based with a pipeline built by their data science team (unsure how specifically). Needs to be finished porting to GCP Dataflow, refine the pipeline, and a bunch of other enhancements.
- Could support new feature development once the backlog is complete.
- Data ingestion coming from 2 different sources:
- A legacy system that is excel file based, parsing and scraping excel files.
- EOS, modern infra with a GCP back end. Have 2 pipelines puling from both sources into GCP for unified source of truth.
- Caremark benefit engine: tool used to make recommendations for how clients could change benefit structures to achieve cost targets. How do changes affect members? How would cost structure change if parameters are updated? Overarching goal is to make it a client-facing product, rolling out 2027 (currently is only internal).
- Legacy SAS code that needs to be migrated into Python and others
- Provide support for modeling work - feature engineering, building datasets for data science teams.
Role Notes:- Roughly 60-70% data pipeline work, 30-40% data curation, data modeling, generation of cleaned datasets for DS teams.
- Broad theme of what are data engineering support needed to make data science team effective: feature engineering that supports model builds.
- Ideal candidate would have CI/CD skills, help put deployment rails in place, becoming more organized and implement best practices around code deployment and sharing/"not deploying on someone's laptop"
- Experience in GCP is important - BigQuery, but would consider "trade-offs" i.e. if a candidate has experience in drug claims and AWS/Azure, they may be able to ramp up more quickly and he would consider that.
- 5+ years of experience - relatively Senior level, self-starter, doesn't need hand-holding.
Team structure:- Team of 5 data scientists split across the 2 projects and 2 data engineers working fractionally across 5 different things. Lack of capacity is driving their gaps in data engineering workflows and pipelines.
- 2 Data Scientists on ACCT: 1 Senior, 1 relatively Junior.
- The Senior data scientist is building the data pipelines (which is out of his wheelhouse), so they are not as performant as they should be.
- Lou is the Architect for this product - not particularly deep on data engineering topics.
- Ideally this Engineer would have perspective on best practices and perspective of how to do things. Would be supported by existing data engineers - ability to point to resources, bounce ideas off of, but ideally they wouldn't be doing code reviews.
- Caremark Benefit Engine is 3 people split across that project - all data science, more technical depth.
- Very parallel, technical folks, data scientists who understand business context, where to find things.
Position Summary:- Supports the creation and operation of Big Data AI/machine learning solutions, using complex healthcare data to produce actionable insights
- Leads and participates in the design, build and management of large-scale data structures and pipelines and efficient Extract/Load/Transform (ETL) workflows.
- Uses strong data analysis skills to profile and validate source datasets to recommend solutions.
- Collaborates with data science team to define and build datasets for training of analytic data models.
- Leverages data management experience to design large scale data structures that support efficient consumption of analytic results.
- Uses strong programming skills in Python, Java or any of the major languages to develop efficient ETL to collect and standardize data to generate insights and addresses reporting needs.
- Develops software to deliver both batch and real time analytics
Required Qualifications:- 5+ years of relevant work experience in Data Engineering/ETL Development
- 5+ years of SQL experience with large datasets
- 2+ years experience with Spark, Python, or Java to build robust data pipelines
- 5+ years of progressively complex related experience.
- Experience leveraging Cloud Technologies (GCP Preferred, AWS, Azure) for ETL and data analysis
- Strong SQL and data analysis experience; data exploration, profiling, and validation.
- Strong collaboration and communication skills within and across teams.
Preferred Qualifications:- Experience with API integration design and development, including authentication, data mapping, and error handling.
- Experience working with Healthcare or Health insurance data
- Understanding of data science methods and statistics
For more information and other jobs available please contact our recruitment team at
careers@tekfortune.com
. To view all the jobs available in the USA and Asia please visit our website at
https://www.tekfortune.com/careers/ .