Platform Engineer

eTeam Inc.

St. Louis, MO

JOB DETAILS
SKILLS
Access Control, Amazon Elastic Compute Cloud (EC2), Amazon Simple Storage Service (S3), Amazon Web Services (AWS), Apache Kafka, Apache Spark, Application Integration, Application Programming Interface (API), Artificial Intelligence (AI), Automation, CPU (Central Processing Unit), Capacity Management, Cloud Computing, Code Reviews, Communication Skills, Computer Science, Continuous Deployment/Delivery, Continuous Integration, DNS (Domain Name System), Data Lake, Data Management, Data Processing, Data Science, Data Storage, Database Design, Database Programming, Database Technology, DevOps, Docker, Electronic Medical Records, Engineering, English Language, Environmental Monitoring, Firewalls, GPU (Graphics Processing Unit), Git, Home Automation, Incident Response, Information Technology & Information Systems, Information/Data Security (InfoSec), Institute of Internal Auditors (IIA), Linux Operating System, Machine Tool, NAT (Network Address Translation), Neo4j, Network Administration/Management, Network Connectivity, Network Monitoring, Network Security, Network Topology, Onboarding, Operational Support, Operations Management, Performance Tuning/Optimization, Production Control, Production Systems, Public Cloud, Python Programming/Scripting Language, REST (Representational State Transfer), Resource Utilization, Sales Management, Scala Programming Language, Scripting (Scripting Languages), Security Compliance, Service Level Agreement (SLA), Software Administration, Software Development, Software Engineering, Splunk, Systems Engineering, Telecommunications Industry, Test Automation, Test Plan/Schedule, Unit Test, Unix Shell Programming, Writing Skills
LOCATION
St. Louis, MO
POSTED
2 days ago
JD:
Role Name: Platform Engineer 
Work site: St. Louis, US (Onsite) Hybrid 3 days a week
Contract

**Platform Engineer IV — IIA (Onshore)**
JOB TITLE: Platform Engineer IV

**JOB SUMMARY**

Infrastructure Intelligence and Analytics (IIA) team builds and operates the data platform and AI agent infrastructure that powers proactive network monitoring and autonomous investigation for network operations. As part of this group, the Platform Engineer IV designs, builds, and maintains the AWS infrastructure that underpins the IIA Data Lake, agent runtime environments, CI/CD pipelines, and graph database systems, and develops the utilitarian application code, automation, and internal tooling for those systems. This role ensures production environments are stable, scalable, and secure while enabling data science and agentic AI workloads to operate reliably at scale.



**MAJOR DUTIES AND RESPONSIBILITIES**

Responsibilities span across the following areas. Individual focus areas will be determined based on team needs and candidate strengths:

**Infrastructure and Data Lake**

· Design and manage AWS infrastructure for the IIA Data Lake including S3 storage, Glue data catalog, Athena query engine, and EMR compute clusters.

· Manage cross-account connectivity, VPC networking, security groups, and IAM roles/policies to enable secure data flow between IIA, upstream data providers, and downstream consumers.

· Build and maintain infrastructure for AI agent runtime environments, including compute resources for LangGraph agents deployed via LangSmith Deployments.

· Support deployment and operation of AWS Neptune for the network topology graph (digital twin), including capacity planning, schema design support, and performance tuning.

· Implement and manage infrastructure-as-code (Terraform, CloudFormation) for repeatable, auditable environment provisioning.

· Manage IAM access key rotations, secrets management (AWS Secrets Manager, Delinea), and security compliance for on-premises and cloud integrations (e.g., Splunk Edge Processor).

**CI/CD and Agent Deployments**

· Build and maintain CI/CD pipelines for AI agent deployments using GitLab CI/CD, Docker, and Artifactory.

· Manage container lifecycle for agents deployed via LangSmith Deployments, including image builds, versioning, and rollback procedures.

· Automate deployment workflows to enable rapid, reliable promotion of agents from development through production.

· Coordinate with SpecGPT platform team on AI Gateway integration, cross-account deployment, and connectivity requirements.

**Application Development and Tooling**

· Develop and maintain utilitarian application code: scripts, CLIs, small services, and automation utilities (primarily Python) that support data ingestion, deployment, environment provisioning, and operational workflows.

· Write integration code and glue services that connect IIA systems with upstream data providers, downstream consumers, and external platforms.

**Production Operations**

· Ensure production environment stability through monitoring, alerting, and incident response. Maintain SLAs for data pipeline availability and agent uptime.

· Implement production monitoring and alerting for deployed agents (health checks, error rates, latency, resource utilization).

· Coordinate with upstream data teams and platform teams (SpecGPT, Splunk, Public Cloud) on connectivity, firewall requests, and integration requirements.

· Support data engineering team with infrastructure needs for new data source onboarding (storage provisioning, access controls, pipeline compute).

· Perform other duties as required.



**REQUIRED QUALIFICATIONS**

**Skills/Abilities and Knowledge**

· Ability to read, write, speak and understand English

· Strong communication skills with ability to explain infrastructure decisions to non-infrastructure stakeholders

· Expert-level experience with AWS services: EC2, S3, IAM, VPC, Glue, Athena, EMR, Secrets Manager, CloudWatch

· Strong experience with infrastructure-as-code (Terraform preferred, CloudFormation acceptable)

· Experience managing cross-account AWS architectures, VPC peering, PrivateLink, and transit gateway configurations

· Experience with IAM policy design, least-privilege access patterns, and service account management

· Experience with containerization (Docker) and container orchestration

· Experience with CI/CD pipelines (GitLab CI preferred)

· Proficiency with Linux-based operating systems and shell scripting

· Experience with monitoring and alerting tools (CloudWatch, Prometheus, Grafana, or similar)

· Understanding of networking fundamentals: DNS, CIDR, NAT, firewalls, security groups

· Demonstrated ability to work across teams and coordinate with external platform owners on connectivity and access requirements

· Proficiency in Python (or a comparable general-purpose language) for building automation, tooling, and applications

· Solid software engineering fundamentals: Git-based workflows, code review, modular and reusable design, dependency management, and writing maintainable, documented code

· Experience writing automated tests (unit/integration) for application and infrastructure code, and integrating those tests into CI/CD

· Ability to write integration code against REST APIs and cloud SDKs (e.g., AWS SDK / boto3)

**PREFERRED QUALIFICATIONS**

**Skills/Abilities and Knowledge**

· Experience with graph databases (AWS Neptune, Neo4j) including deployment, scaling, and operational management

· Experience with Apache Kafka or similar streaming platforms

· Experience with Apache Spark (Scala preferred) for distributed data processing

· Experience with Airflow or similar workflow orchestration platforms

· Experience in the telecommunications industry or other large-scale network operations environments

· Familiarity with AI/ML infrastructure requirements (model serving, GPU/CPU compute, artifact management via MLflow or similar)

· Experience with Splunk integration, particularly Edge Processor and ClientP connectivity

· AWS certifications (Solutions Architect, DevOps Engineer, or similar)

· Experience developing and operating small services or APIs (e.g., FastAPI/Flask) in a production environment

**Education**

Bachelor's degree in Computer Science, Information Technology, Systems Engineering, or related field, or relevant experience



**Related Experience**

· Bachelor's degree: 5 years of platform/infrastructure engineering experience

· Master's degree: 3 years of platform/infrastructure engineering experience

**WORKING CONDITIONS**

Hybrid (3 days in office and 2 days remote)
 

About the Company

e

eTeam Inc.

Looking for a great job? Join eTeam. We’re looking for talented staffing professionals to join our staff. We also provide contract assignments and full-time jobs at Fortune 2000 Companies. We’ve been named one of the best companies to work for by Staffing Industry Analysts and New Jersey Business.
COMPANY SIZE
100 to 499 employees
INDUSTRY
Other/Not Classified
FOUNDED
1998
WEBSITE
www.eteaminc.com