AI Engineering #deepeval#ragas#evaluation#testing#ci-cd#agents#rag#github-actions

Evaluation-Driven Agent Development: Automated Regression Testing with DeepEval and Ragas

S

S L Manikanta

Sep 13, 2026 • 6 min read

bolt Key Takeaways

  • DeepEval provides LLM-as-judge metrics (faithfulness, answer relevancy, hallucination) that you can run as pytest test cases with assertion-style pass/fail gates.
  • Ragas specializes in RAG pipeline evaluation: context_precision, context_recall, and answer_correctness against a golden Q&A dataset.
  • Run evals in GitHub Actions on every PR that changes prompts, agent logic, or model versions — block merges if metrics drop below defined thresholds.
  • Store your evaluation dataset in a versioned JSONL file in the repo — not in the eval framework's cloud — so dataset changes are code-reviewed like any other change.
✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

[!NOTE] Quick Start — DeepEval in pytest (2 minutes):

pip install deepeval ragas
# test_agent.py
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import FaithfulnessMetric, AnswerRelevancyMetric

def test_rag_faithfulness():
    test_case = LLMTestCase(
        input="What is LangGraph's checkpointing mechanism?",
        actual_output=your_rag_pipeline("What is LangGraph's checkpointing mechanism?"),
        retrieval_context=["LangGraph uses a checkpointer to persist state to Postgres..."],
    )
    assert_test(test_case, [FaithfulnessMetric(threshold=0.7)])

Run: pytest test_agent.py -v

Most agent teams realize they need regression tests when a prompt change silently broke production quality three deploys ago. Evaluation-driven development builds the test harness before that happens — so prompt changes, model upgrades, and RAG pipeline edits are gated by measured quality, not vibes.


Environment

PackageVersion
deepeval1.4.0
ragas0.1.21
pytest8.2+
datasets2.20+
Python3.11+
Judge LLMClaude 3.5 Haiku or GPT-4o-mini

1. Evaluation Architecture

graph TD
    PR[GitHub Pull Request\n prompt / model / RAG change]
    PR --> CI[GitHub Actions: eval.yml]
    CI --> Build[Build Agent / RAG Pipeline]
    Build --> Load[Load Evaluation Dataset\n eval_dataset.jsonl]
    Load --> Run[Run Pipeline on Each Sample]
    Run --> Metrics[DeepEval + Ragas Metrics\n Faithfulness, Relevancy, Recall]
    Metrics --> Gate{All metrics\nabove threshold?}
    Gate -- Yes --> Merge[Allow Merge]
    Gate -- No --> Block[Block Merge\n Post metric delta as PR comment]

2. Evaluation Dataset Format

Store your dataset as versioned JSONL in the repo. Each row is a test case:

{"id": "q001", "input": "What is LangGraph's checkpointing mechanism?", "expected_output": "LangGraph uses a checkpointer to persist graph state to a durable store like Postgres after each node executes.", "context": ["LangGraph supports pluggable checkpointers. The AsyncPostgresSaver writes the full agent state to a Postgres database after every node completes."]}
{"id": "q002", "input": "How do I fix vLLM CUDA OOM errors?", "expected_output": "Reduce max-model-len to cap KV cache allocation and enable chunked prefill.", "context": ["vLLM CUDA OOM is typically caused by KV cache over-allocation. Setting --max-model-len to 4096 reduces the reservation significantly."]}

Load it in your tests:

import json
from pathlib import Path

def load_eval_dataset(path: str = "tests/eval_dataset.jsonl") -> list[dict]:
    return [json.loads(line) for line in Path(path).read_text().splitlines() if line.strip()]

3. DeepEval: Agent and RAG Metrics

Core Metrics

from deepeval.metrics import (
    FaithfulnessMetric,      # Does the answer contradict the retrieved context?
    AnswerRelevancyMetric,   # Does the answer address the question?
    ContextualPrecisionMetric, # Are retrieved chunks relevant to the question?
    HallucinationMetric,     # Does the output contain facts not in the context?
)
from deepeval import assert_test
from deepeval.test_case import LLMTestCase

faithfulness = FaithfulnessMetric(threshold=0.70, model="gpt-4o-mini")
relevancy = AnswerRelevancyMetric(threshold=0.75, model="gpt-4o-mini")

Test Suite

import pytest
from your_agent import run_rag_pipeline  # your actual pipeline

@pytest.mark.parametrize("sample", load_eval_dataset())
def test_rag_quality(sample: dict):
    # Run your RAG pipeline
    output, retrieved_chunks = run_rag_pipeline(sample["input"])

    test_case = LLMTestCase(
        input=sample["input"],
        actual_output=output,
        expected_output=sample["expected_output"],
        retrieval_context=retrieved_chunks,
    )

    assert_test(test_case, metrics=[
        FaithfulnessMetric(threshold=0.70, model="gpt-4o-mini"),
        AnswerRelevancyMetric(threshold=0.75, model="gpt-4o-mini"),
    ])

4. Ragas: RAG-Specific Pipeline Evaluation

Ragas evaluates the full RAG pipeline — retriever quality and generator quality together:

from ragas import evaluate
from ragas.metrics import (
    faithfulness,        # Generated answer vs retrieved context
    answer_relevancy,   # Generated answer vs question
    context_precision,  # Are retrieved chunks ranked correctly?
    context_recall,     # Do retrieved chunks cover the expected answer?
    answer_correctness, # Generated answer vs ground truth
)
from datasets import Dataset

def build_ragas_dataset(eval_samples: list[dict], pipeline_fn) -> Dataset:
    """Run the pipeline on each sample and build a Ragas-compatible Dataset."""
    rows = []
    for sample in eval_samples:
        output, contexts = pipeline_fn(sample["input"])
        rows.append({
            "question": sample["input"],
            "answer": output,
            "contexts": contexts,
            "ground_truth": sample["expected_output"],
        })
    return Dataset.from_list(rows)

def run_ragas_evaluation(eval_samples: list[dict], pipeline_fn) -> dict:
    dataset = build_ragas_dataset(eval_samples, pipeline_fn)
    result = evaluate(
        dataset,
        metrics=[faithfulness, answer_relevancy, context_precision, context_recall],
    )
    return result

5. Threshold Enforcement

# tests/test_ragas_regression.py
import pytest
from your_agent import run_rag_pipeline

THRESHOLDS = {
    "faithfulness": 0.70,
    "answer_relevancy": 0.75,
    "context_precision": 0.65,
    "context_recall": 0.60,
}

def test_ragas_regression():
    samples = load_eval_dataset()
    scores = run_ragas_evaluation(samples, run_rag_pipeline)

    failures = []
    for metric, threshold in THRESHOLDS.items():
        actual = scores[metric]
        if actual < threshold:
            failures.append(
                f"{metric}: {actual:.3f} < threshold {threshold:.3f} (delta: {actual - threshold:+.3f})"
            )

    if failures:
        fail_msg = "Evaluation regression detected:\n" + "\n".join(failures)
        pytest.fail(fail_msg)
    else:
        for metric, threshold in THRESHOLDS.items():
            print(f"  {metric}: {scores[metric]:.3f} (threshold: {threshold:.3f}) PASS")

6. GitHub Actions Integration

# .github/workflows/eval.yml
name: Agent Evaluation Gate

on:
  pull_request:
    paths:
      - "src/agent/**"
      - "src/rag/**"
      - "prompts/**"
      - "tests/eval_dataset.jsonl"

jobs:
  evaluate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"

      - name: Install dependencies
        run: pip install deepeval ragas pytest

      - name: Run evaluation gate
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: pytest tests/test_ragas_regression.py tests/test_agent.py -v --tb=short

      - name: Post results as PR comment
        if: always()
        uses: actions/github-script@v7
        with:
          script: |
            const fs = require('fs');
            // Read pytest output from a file if saved, or use inline reporting
            github.rest.issues.createComment({
              issue_number: context.issue.number,
              owner: context.repo.owner,
              repo: context.repo.repo,
              body: "Evaluation gate results posted. Check the Actions log for metric breakdown."
            });

7. Metric Interpretation Guide

MetricWhat It MeasuresLow Score Means
faithfulnessDoes the answer contradict the retrieved context?LLM is hallucinating beyond the retrieved chunks
answer_relevancyDoes the answer address the question asked?LLM is generating off-topic content
context_precisionAre retrieved chunks ranked by relevance?Retriever is returning irrelevant chunks high in results
context_recallDo retrieved chunks contain the expected answer?Retriever is missing key documents
answer_correctnessIs the answer factually correct vs ground truth?Full pipeline has factual errors
hallucination (DeepEval)Does output contain facts not in any context?Model is fabricating information

8. Cost Management

StrategySavings
Use gpt-4o-mini or claude-3-5-haiku as judge (not GPT-4o)80-90% cost reduction vs frontier model judge
Run evals only on path-filtered PRsEliminates evals on unrelated changes
Cache judge responses by input hashEliminates redundant re-scoring of unchanged test cases
Keep eval dataset under 100 samples100 samples at $0.003/call = $0.30/run
Batch judge API calls where supportedReduces latency; some APIs offer batch discounts

Next Steps

For structured output evaluation specifically (checking JSON schema compliance), see Enforcing Deterministic JSON Schemas with Outlines and SGLang.

For the RAG pipeline that you’re evaluating, see Local RAG with Ollama, nomic-embed-text, and LanceDB.

For LangGraph agent architecture that the eval suite gates, see Building Resilient LangGraph Workflows with Async Postgres Checkpointing.

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.

Related Articles

AI Engineering
Advanced RAG on Azure: Hybrid Search & Re-ranking Implementation

Going beyond basic vector search. A technical guide to implementing Hybrid Search (Keyword + Vector) and Semantic Re-ranking using Azure AI Search and OpenAI.

AI Engineering
The Shift to Agentic AI: Why Enterprise Architecture is Moving Beyond Chatbots

Chatbots are dead. Welcome to the era of Agentic AI. Explore how enterprises are deploying autonomous agents for complex workflows, the architectural shift required, and the rise of specialized inference models like Nemotron 3.5 Lightning.

AI Engineering
Mastering Agent Skills: A New Standard for AI Capabilities

An in-depth guide on Agent Skills, exploring how to extend AI agents like Claude with specialized knowledge, workflows, and tools using an open, filesystem-based format.