Sr. IT Systems Engineer- Azure DCM Infotech
- Full-time
| Location | Georgia, GA |
Back to Search
Senior ML / Evaluation Engineer
Remote in Georgia, & 5 others
Python.AI& 5 others
apply
FacebookLinkedInSend via email
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a job
Location-specific conditions & benefits*
Choose an option
Join the Enterprise Agent Development Platform project at EPAM. We are building a cloud-native platform that enables engineering teams to develop, deploy, and operate AI agents in production.
The platform combines modern agent frameworks, AWS infrastructure, CI/CD, observability, and governance to make AI development faster and more reliable.
As a Python AI Evaluation Engineer, you will own a key part of this platform: building the capabilities that measure and validate the quality of AI solutions. You will design evaluation approaches, develop custom evaluators, and integrate quality checks into the software delivery lifecycle.
This role is a strong fit for an engineer who enjoys solving new GenAI challenges and turning them into practical, automated engineering solutions.
Responsibilities
Design and implement evaluation frameworks for LLM-based applications and AI agents
Develop LLM-as-a-Judge and deterministic, code-based evaluators
Build custom Python evaluators for quality and behavioral checks
Define evaluation criteria, metrics, thresholds, and acceptance rules
Evaluate agent behavior across individual responses, tool calls, and complete workflows
Work with OpenTelemetry traces and spans as evaluation data
Integrate evaluations into CI/CD pipelines and automated deployment gates
Enable continuous quality monitoring of solutions in production
Establish reusable evaluation patterns and engineering standards
Work closely with AI Engineers, Architects, and Platform Engineers to embed quality into the development process
Requirements
5+ years of experience in ML Engineering, AI Engineering, or AI Platform Engineering
Strong Python development experience
Hands-on experience with LLM/GenAI evaluation
Experience designing and implementing evaluation frameworks
Experience developing custom or deterministic evaluators
Experience integrating AI/ML quality checks into CI/CD
Good understanding of LLM and AI agent architectures
Nice to have
Hands-on experience with AWS AgentCore Evaluation
Experience with AWS Bedrock Guardrails, including PII detection
Knowledge of CloudWatch metrics and production monitoring
Experience with OpenTelemetry
Familiarity with LangGraph, Strands Agents, or similar agent frameworks




