Developer Productivity #ci-cd#ai agents#evals#testing#azure devops#github actions

CI/CD Pipelines for AI Agents: Automated Evaluation and Regression Testing

S

S L Manikanta

Aug 28, 2026 • 5 min read

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

When software engineers modify code, unit tests and integration tests verify that nothing broke. But when you tweak a system prompt, update a tool schema, or switch underlying model providers, traditional unit tests usually fail to catch behavioral regressions.

An updated prompt might fix a customer edge case while secretly degrading your agent’s SQL generation accuracy or causing subtle hallucinations.

To ship AI agents with confidence, you need automated evaluation pipelines (Evals) integrated directly into your continuous integration workflow.

Here is how to set up automated agent evaluation pipelines in GitHub Actions and Azure DevOps to block breaking changes on pull requests.

flowchart TD
    Dev[Developer Opens Pull Request] --> Trigger[Trigger CI Pipeline]
    
    subgraph CI [CI / CD Test Runner]
        Setup[Install Dependencies & Agent Code] --> UnitTests[Run Fast Deterministic Unit Tests]
        UnitTests --> EvalSuite[Run Agent Evaluation Suite: DeepEval / Promptfoo]
        
        EvalSuite --> Metric1[Hallucination Metric: Threshold > 0.90]
        EvalSuite --> Metric2[Tool Calling Accuracy: Threshold > 0.95]
        EvalSuite --> Metric3[Answer Relevance: Threshold > 0.85]
    end

    Metric1 --> Check{Pass All Quality Gates?}
    Metric2 --> Check
    Metric3 --> Check

    Check -->|Pass| Merge[PR Approved: Auto Merge]
    Check -->|Fail| Block[PR Blocked: Post Detailed Regression Report to PR]

The Three Types of Agent Tests in CI

An effective evaluation pipeline runs three layers of testing:

  1. Deterministic Unit Tests: Mock LLM responses with canned JSON payloads to verify that tool parsing, state machines, and data validation logic work correctly. These are cheap, fast (under 10 seconds), and run on every commit.
  2. Golden Dataset Regression Tests: Run the agent against a curated set of 20 to 50 realistic user queries where the expected tool calls and ground truth answers are known.
  3. LLM as a Judge Semantic Metrics: Use a fast, judge model (like GPT-4o-mini or Claude 3.5 Haiku) to grade answers on specific criteria like answer relevancy, faithfulness, and absence of hallucination.

Writing an Evaluation Test Suite with DeepEval

DeepEval is an open source evaluation framework for Python that integrates directly with standard pytest.

1. Installation

pip install deepeval pytest

2. Test Implementation (test_agent_evals.py)

Here is an automated test suite that validates an AI customer support agent against hallucination and answer relevancy metrics:

import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import HallucinationMetric, AnswerRelevancyMetric, GEval
from deepeval.test_case import LLMTestCaseParams

# Simulated agent invocation function
def invoke_agent(user_input: str) -> str:
    # In a real setup, call your LangGraph or PydanticAI agent here
    return "Our standard shipping takes 3 to 5 business days. Expedited shipping takes 1 to 2 business days."

# Define test cases with retrieval context and expected outputs
TEST_CASES = [
    {
        "input": "How long does standard delivery take?",
        "context": ["Shipping Policy: Standard shipping delivers in 3-5 business days across the US. Overnight shipping takes 1 day."],
        "expected_topics": ["3 to 5 business days"]
    },
    {
        "input": "Can I return an opened software license?",
        "context": ["Return Policy: Opened software licenses are non-refundable once activated."],
        "expected_topics": ["non-refundable"]
    }
]

@pytest.mark.parametrize("test_data", TEST_CASES)
def test_agent_hallucination_and_relevancy(test_data):
    actual_output = invoke_agent(test_data["input"])
    
    test_case = LLMTestCase(
        input=test_data["input"],
        actual_output=actual_output,
        context=test_data["context"]
    )
    
    # Metric 1: Check that the agent did not invent facts outside the context
    hallucination_metric = HallucinationMetric(threshold=0.8)
    
    # Metric 2: Check that the answer directly addresses the user question
    relevancy_metric = AnswerRelevancyMetric(threshold=0.8)
    
    # Assert both metrics pass quality thresholds
    assert_test(test_case, [hallucination_metric, relevancy_metric])

When you run pytest test_agent_evals.py, it calculates scores from 0.0 to 1.0. If any score drops below the defined threshold (e.g. 0.8), the test fails with a clear explanation of why the output was flagged.


Setting Up the CI/CD Pipeline

1. GitHub Actions Workflow (.github/workflows/agent-evals.yml)

name: AI Agent Evaluation Suite

on:
  pull_request:
    branches: [main]
    paths:
      - 'src/agents/**'
      - 'prompts/**'
      - 'tests/evals/**'

jobs:
  run-evals:
    runs-on: ubuntu-latest
    steps:
      - name: Checkout Code
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: '3.11'
          cache: 'pip'

      - name: Install Dependencies
        run: |
          pip install -r requirements.txt
          pip install deepeval pytest

      - name: Run Deterministic Unit Tests
        run: pytest tests/unit/

      - name: Run LLM Evaluation Quality Gates
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
        run: |
          # Run evals and output JUnit XML report
          deepeval test run tests/evals/test_agent_evals.py

2. Azure DevOps Pipeline (azure-pipelines-evals.yml)

trigger: none
pr:
  branches:
    include:
      - main
  paths:
    include:
      - src/agents/*
      - prompts/*

pool:
  vmImage: 'ubuntu-latest'

variables:
  - group: ai-pipeline-secrets # Contains OPENAI_API_KEY

steps:
- task: UsePythonVersion@0
  inputs:
    versionSpec: '3.11'

- script: |
    python -m pip install --upgrade pip
    pip install -r requirements.txt
    pip install deepeval pytest
  displayName: 'Install Dependencies'

- script: |
    pytest tests/evals/test_agent_evals.py --junitxml=eval-results.xml
  env:
    OPENAI_API_KEY: $(OPENAI_API_KEY)
  displayName: 'Execute Agent Evals'

- task: PublishTestResults@2
  condition: always()
  inputs:
    testResultsFormat: 'JUnit'
    testResultsFiles: '**/eval-results.xml'
    testRunTitle: 'AI Agent Evaluation Results'

Best Practices for Cost and Speed in CI

  1. Test on Golden Subsets in PRs: Running 500 test cases on every single commit gets expensive and slows down development. Run a fast 20 question subset on pull requests, and run the full 500 question benchmark nightly on a schedule.
  2. Cache LLM Responses for Unchanged Nodes: If a developer only changes the prompt for Tool A, avoid re-evaluating Tool B. Use caching layers to keep eval runs fast.
  3. Use Cheap Models as Judges: Use cost efficient models like gpt-4o-mini or claude-3-5-haiku as evaluation judges rather than expensive flagship models.
✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.

Related Articles

Developer Productivity
Infrastructure as Code: Getting Started with Terraform in Azure DevOps

Automate your cloud infrastructure securely. A step-by-step guide to integrating Terraform with Azure Pipelines, managing state, and handling secrets.

Developer Productivity
The Architect's Guide to Azure Pipelines: Patterns for Enterprise CI/CD

Going beyond Hello World: A deep dive into scalable YAML templates, security governance, dynamic environments, and cost-optimization strategies for Azure DevOps.

Developer Productivity
Docker in Production: Optimization and Security for Kubernetes

A comprehensive guide to building slim, secure, and fast container images. From multi-stage builds to rootless containers and health probes.