Softwar Engineer AI Evaluation & Automation

Tranzeal Inc.

  • San Diego, CA
  • 23 days ago

    Highlights

    Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time. Key Responsibilities " Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.

    Numbers & Facts

    LocationSan Diego, CA

    Description

    Job Title: Software Engineer AI Evaluation & Automation
    Location: San Diego, CA
    Job Description:
     Role Overview
    
    Help build and scale the tooling we use to measure how well AI-powered software
    development tools actually perform. You'll develop evaluation harnesses, automate
    benchmark runs, and help make sure the results we produce are reproducible and hold up
    to scrutiny. This is an engineering role, but a lot of the work is about getting the
    measurement right, not just automating it.
    
    Key Responsibilities
    " Build and integrate evaluation harnesses and automation for software development use
    cases, including turning real engineering artifacts like merged pull requests into
    repeatable benchmark tasks.
    " Build versioned, repeatable processes to evaluate AI tools, models, and harnesses,
    with reproducible run environments (pinned dependencies, containerized runs,
    isolated worktrees) so results stay comparable over time.
    " Validate and calibrate evaluation approaches against human judgment, so scores are
    consistent and correct rather than just repeatable.
    " Support execution-based benchmarking across quality, productivity, and eJiciency
    measures, including cost and latency.
    " Analyze results across repeated runs, looking at variance, failure patterns, and cost per
    outcome, and find ways to make the workflows more reliable and more automated.
    " Work with engineering and data teams to improve the tooling, and document how the
    evaluations work and what they found for both technical and leadership audiences.
    Required Skills & Experience
    " Strong software engineering background, with real experience building automation,
    developer tooling, or test and validation systems.
    " Proficient in at least one general-purpose language such as Python, Java, or JavaScript
     the specific language background is flexible.
    " Solid working knowledge of Git, including how branches, history, and working trees
    behave, and of containerization with Docker.
    " Experience with APIs, development environments, CI/CD pipelines, and typical
    engineering workflows.
    " Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such
    as why a judge can be consistent but still wrong, why a single run can mislead, and how
    benchmark contamination happens.
    " Able to troubleshoot technical problems, think clearly about whether a measurement is
    valid, and analyze results carefully.
    " Hands-on experience using AI coding tools and agentic harnesses such as Claude
    Code, Devin, or Cursor, and command of the best practices for working with them
    effectively.
    
    Preferred Experience
    " Experience designing benchmarks or evaluations for software systems, especially
    execution-based grading that verifies against tests.
    " Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them
    against human raters.
    " Experience with build-system-aware test selection, such as Bazel or mapping changed
    files to the tests that cover them.
    " Experience building reproducible test environments and managing versioned
    evaluation datasets.
    " Comfortable writing up methodology and results for engineering leadership. 

    Similar Jobs

    See more jobs