We operate a multi-tenant automotive SaaS platform serving thousands of dealer groups across the United States. Our backend - event-driven serverless on AWS - orchestrates everything from dealer onboarding to inventory management to real-time transaction processing. That platform works. Now we need to make it think - and keep it modern, secure, and safe to ship. We also run a monolith that needs upkeep while we build the next-gen platform and redesign its components, a few at a time, into microfrontends.
This is a hands-on engineering role at its core. Day to day you build, ship, and operate agentic AI systems - autonomous, tool-using agents (AgentCore, MCP servers) that observe platform state, reason over dealer context, and take action through production APIs. But you are also the engineer who keeps the whole platform current and trustworthy: you continuously hunt and fix security vulnerabilities, upgrade the frameworks and runtimes underneath us, drive adoption of an agentic development harness to speed up delivery, and lead modernization of our existing systems.
You work across the org, not inside one team"s walls. You own production monitoring and the SRE practices that keep us reliable, and you are the champion for safe releases - defining the checks and balances that gate production, setting the engineering standards, and making sure they are actually enforced. When something misbehaves in production, your instrumentation, your guardrails, and your release discipline are what make it fail safe instead of fail loud.
Reports to: SVP, Engineering.
What You"ll Own
- Hands-on agentic development - designing, building, and operating agentic AI systems (AWS Bedrock AgentCore, MCP servers) every day: agent code, tool interfaces, evaluation harnesses, and production AI workflows.
- Driving adoption of the agentic development harness - making agent-assisted development a first-class way the org ships, and measurably improving speed of delivery across teams.
- Security vulnerability management - continuously monitoring for vulnerabilities (dependencies, CVEs, container images, IAM/config drift), remediating quickly, and keeping frameworks, libraries, and runtimes patched and upgraded.
- Platform modernization - creating and driving initiatives to modernize existing platforms: retiring legacy patterns, adopting current frameworks, and staying close to the leading edge.
- Production monitoring & SRE - owning observability (OpenTelemetry, CloudWatch), SLOs and error budgets, on-call and incident response.
- Third-party integration observability - a standardized way to organize all third-party integration calls, report on them with precision, and run anomaly detection on call volumes with continuous monitoring, so a spike, drop, or failure surfaces before it becomes an outage or a cost problem.
- On-call cookbook - keeping applications well documented so an on-call engineer can quickly query a cookbook, understand the likely problem areas, and know where to look.
- Release governance - champion across releases: defining and enforcing the checks and balances (CI/CD quality gates, security and test gates, rollback criteria, progressive delivery) required to move change into production.
- Setting and enforcing engineering standards - establishing the patterns, guardrails, and review bar the org follows, and holding the line so they are consistently applied.
- Working across the org - partnering with every engineering team to raise the technical, security, and reliability bar broadly.
Tech & Tools
- Cloud: Lambda, EventBridge, DynamoDB, S3, ECS Fargate, Aurora, API Gateway, CloudWatch, Secrets Manager.
- AI & Agentic: AWS Bedrock AgentCore, MCP servers, LangChain/LangGraph.
- Languages: Python and Java (Spring Boot), TypeScript/React (Next.js), legacy PHP/Laravel.
- Security: Dependabot/Snyk, SAST/DAST, container image scanning, secrets management, IAM hardening.
- Reliability & Observability: OpenTelemetry, CloudWatch, SLO/error-budget tooling, PagerDuty, Datadog.
- CI/CD & Release: CircleCI, CloudFormation, progressive delivery, quality/security gates, rollback automation.
- Integration Surfaces: REST, SOAP/XML, EventBridge, SES, Playwright.
How You"ll Use AI
This is not a "we are AI-curious" company. This is the role that makes agent-assisted development the norm for everyone else - building and operating production agents, triaging CVEs and generating patch/upgrade PRs, driving framework upgrades, reviewing PRs across the stack, standing up observability, and drafting ADRs and runbooks.
Hands-On Expectations
Roughly 60-70% building and operating, 20-30% on design and standards, and ~10% on cross-team enablement.
First 12 Months
- Months 1-3: Immerse in the codebase, ship your first meaningful changes, audit our security posture, dependency/framework currency, and release pipeline, and publish a baseline of engineering standards and release checks.
- Months 4-6: Ship agentic automation into at least one production workflow, roll out the agentic development harness, stand up the SRE baseline (SLOs, on-call, dashboards), and automate vulnerability scanning and patching.
- Months 7-9: Drive a modernization initiative end to end; harden the release gates and enforce them across teams.
- Months 10-12: Measurable delivery-speed gains from agentic adoption, production reliability against SLOs, and engineering standards adopted org-wide.