Cloud FinOps and AI Platform Engineer

CyberData Technologies

  • Bethesda, MD
  • 1 day ago

    Highlights

    The Principal Multi-Cloud Architect leads architecture, security, and federal compliance; the Senior Cloud Platform Engineer leads cloud engineering, Infrastructure as Code, CI/CD, automation, and platform reliability; and this role leads FinOps/Kion and supports AI-platform operations. This role has primary ownership of Kion-based multi-cloud financial operations and AI cost management, with adjacent responsibility for operational support of NHGRI's secure institutional AI platform using LibreChat, LiteLLM, AWS Bedrock, Azure OpenAI, and GCP Vertex AI.

    Numbers & Facts

    LocationBethesda, MD

    Description

    Position Summary

    CyberData Technologies is seeking a Cloud FinOps and AI Platform Engineer to support a complex federal research environment spanning Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP).

    This role has primary ownership of Kion-based multi-cloud financial operations and AI cost management, with adjacent responsibility for operational support of NHGRI's secure institutional AI platform using LibreChat, LiteLLM, AWS Bedrock, Azure OpenAI, and GCP Vertex AI. The engineer will convert cloud and AI usage data into actionable controls, forecasts, optimization recommendations, and reliable platform operations for technical, research, security, and executive stakeholders.

    The position works as part of an integrated technical team. The Principal Multi-Cloud Architect leads architecture, security, and federal compliance; the Senior Cloud Platform Engineer leads cloud engineering, Infrastructure as Code, CI/CD, automation, and platform reliability; and this role leads FinOps/Kion and supports AI-platform operations. This is not a data-science or model-development position.

    Primary Responsibilities

    Kion, FinOps, and Cloud Optimization

    Manage and operate Kion as the centralized spend-governance platform across AWS, Azure, and GCP.

    Maintain integration of cloud accounts, subscriptions, billing accounts, organizational units, projects, and funding sources into Kion and related reporting workflows.

    Configure and maintain monthly and full-budget alerts, passive spend enforcement, and approved active enforcement mechanisms in coordination with NHGRI stakeholders.

    Maintain account-level and service-level cost allocation, tagging, showback/chargeback, forecasting, anomaly review, and budget-control practices.

    Maintain a current inventory of cloud accounts, projects, subscriptions, credits, and research-billing arrangements.

    Prepare biweekly Kion spend metrics and monthly cost, usage, risk, and optimization inputs for management and security reporting.

    Analyze consumption patterns, identify waste and right-sizing opportunities, evaluate provider and service tradeoffs, and track recommended actions through implementation and validation.

    Support the initial and semiannual Cloud Optimization Assessments, including cost baselines, resource-efficiency findings, prioritized actions, and monthly progress snapshots.

    Support Principal Investigators and research laboratories with Terra billing projects, cloud credits, grant and laboratory funding, and related coordination across AnVIL, the All of Us Research Program Workbench, STRIDES, and NIH stakeholders.

    AI Platform Operations and AI Cost Management

    Operate and support a secure multi-provider AI platform using LibreChat, LiteLLM, or comparable gateway and user-interface technologies.

    Support integrations with AWS Bedrock, Azure OpenAI, and GCP Vertex AI, including model endpoint changes, provider routing, version updates, quotas, rate limits, and service-availability monitoring.

    Implement token-level usage tracking and cost attribution by user, model, provider, organizational unit, or other approved allocation dimension.

    Integrate AI-service consumption with Kion and related reporting to provide near-real-time visibility against Institute budgets.

    Perform token-cost analysis, budget forecasting, model and provider cost-benefit comparisons, and budget-capping or alerting configuration for AI experimentation and production use.

    Partner with the Principal Multi-Cloud Architect and Senior Cloud Platform Engineer to implement approved identity, least-privilege access, private connectivity, data-residency, logging, and monitoring controls.

    Integrate AI usage and operational telemetry with Splunk, OpenSearch, Kion, and cloud-native monitoring services for security auditing, troubleshooting, and operational reporting.

    Develop and maintain runbooks, SOPs, knowledge articles, dashboards, and operational documentation for platform access, provider changes, incidents, cost controls, and recovery activities.

    Support researchers, Principal Investigators, ITB personnel, and leadership with approved AI-service onboarding, use-case evaluation, cost considerations, and knowledge-sharing sessions.

    Contribute technical evidence and implementation input to AI governance, NIST AI Risk Management Framework activities, security authorization documentation, and governed agentic-AI workflows.

    Team and Operational Responsibilities

    Coordinate with NIH CIT, identity, network, subscription, security, research-program, and platform-provider teams to resolve dependencies and maintain service continuity.

    Troubleshoot FinOps, Kion, cloud-billing, AI-gateway, integration, performance, and service-availability issues and drive corrective actions to closure.

    Participate in transition, knowledge-transfer, workload-planning, quality-control, and cross-training activities with the broader delivery team.

    Provide limited after-hours support for maintenance, incidents, cloud-service disruptions, or platform-recovery activities.

    Required Qualifications

    Bachelor's degree in computer science, information technology, engineering, data analytics, finance, business, or a related field, or equivalent relevant experience.

    Five or more years of relevant experience in cloud engineering, cloud operations, FinOps, platform operations, DevOps, systems engineering, or a closely related discipline.

    At least three years of hands-on experience with cloud financial management, cost optimization, cloud governance, or production cloud-platform operations.

    Practical experience with cloud billing and cost-management capabilities in at least two major cloud providers - AWS, Azure, or GCP - with working knowledge of the third.

    Experience using Kion or a comparable multi-cloud FinOps, governance, or cloud financial-management platform. Direct Kion experience is strongly preferred.

    Demonstrated experience with budgets, forecasts, tagging, allocation, showback/chargeback, spend alerts, enforcement controls, anomaly review, and optimization recommendations.

    Experience operating or supporting an AI platform, API gateway, developer platform, or other multi-tenant cloud service, with exposure to at least one of AWS Bedrock, Azure OpenAI, or GCP Vertex AI.

    Experience with scripting, APIs, automation, or Infrastructure as Code using technologies such as Python, PowerShell, Terraform, GitHub, serverless functions, or workflow tools.

    Working knowledge of cloud identity, private endpoints, networking, logging, monitoring, and least-privilege access controls.

    Ability to analyze technical and financial data and present clear recommendations to technical, research, business, security, and executive stakeholders.

    Strong troubleshooting, technical-writing, documentation, and communication skills.

    Ability to obtain and maintain required HHS/NIH suitability, badging, training, and system access.

    Preferred Qualifications

    Direct experience administering Kion, including multi-cloud account integration, budget alerting, spend enforcement, reporting, and configuration-compliance capabilities.

    Experience with LibreChat, LiteLLM, or another multi-LLM gateway or model-routing platform.

    Experience managing AI model endpoints, provider routing, API access, quotas, rate limits, model versions, usage telemetry, or token-level cost attribution.

    Experience integrating cloud or AI telemetry with Splunk, OpenSearch, or comparable observability and security platforms.

    Experience with Terraform, GitHub Enterprise Cloud, Cloud Custodian, AWS Lambda, Azure Functions, or comparable automation technologies.

    Familiarity with Microsoft Entra ID, Azure Privileged Identity Management, AWS PrivateLink, Azure ExpressRoute, AWS IAM, or GCP IAM.

    Familiarity with the NIST AI Risk Management Framework, FISMA, HHS/NIH policy, security assessment and authorization, or other regulated federal-cloud requirements.

    Experience supporting scientific, biomedical, genomic, research-computing, or other data-intensive mission environments.

    Familiarity with Terra billing projects, NIH STRIDES, AnVIL, NIH Data Commons, the All of Us Research Program Workbench, or similar research-cloud ecosystems.

    Experience facilitating knowledge-sharing sessions, developing SOPs, and supporting technical communities of practice.

    Preferred Certifications

    One or more current certifications in FinOps, cloud operations, AI platforms, automation, or cloud security are desired. Relevant certifications include:

    FinOps Certified Practitioner or FinOps Certified Professional

    AWS Certified Solutions Architect - Associate, AWS Certified CloudOps Engineer - Associate, or AWS Certified AI Practitioner

    Microsoft Certified: Azure Administrator Associate, Azure AI Fundamentals, or another current Azure AI certification

    Google Associate Cloud Engineer or Google Professional Machine Learning Engineer

    HashiCorp Certified: Terraform Associate

    Splunk Core Certified Power User, Splunk Cloud Certified Admin, or related Splunk certification

    Other relevant FinOps, cloud-governance, AI-platform, DevOps, cloud-operations, or cloud-security certification

     


    Similar Jobs

    See more jobs