RAG + RAGAS: Turning AI Answer Quality into Numbers
AI Engineering·10 min

RAG + RAGAS: Turning AI Answer Quality into Numbers

Building a RAG pipeline for education AI, then using RAGAS to score quality on four metrics. The traps I hit, and why low Faithfulness can mislead.

Y
Young

Recently I went deep on RAG in an education AI project and wired RAGAS evaluation directly into the pipeline

Here's what I learned, including some non-obvious parts


What RAG Is, and Why Education Needs It

RAG (Retrieval Augmented Generation) does one simple thing:

Before answering a question, search a knowledge base for relevant passages, then give those passages plus the question to the model

Why not just ask the model directly?

Models have a training cutoff and can't access private knowledge. In education this matters even more — textbooks, curricula, and worked solutions need version control. You can't have the model hallucinate answers


Stack Selection

The RAG stack for this project:

LayerTechnologyWhy
Vector DBpgvectorNative PostgreSQL extension — no separate vector DB to operate
EmbeddingVertex AI text-embedding-004GCP ADC auth, no extra API key; 768 dimensions
RetrieverLangChain BaseRetrieverNative integration with LangGraph + RAGAS
LLMLiteLLM wrapperOne-line model swap; supports Google, OpenAI, open source
EvaluationRAGASIndustry standard, four core metrics

pgvector vs a standalone vector database

Many people's first instinct is Pinecone or Weaviate, but for this project there was no reason to operate an extra service. pgvector runs on top of PostgreSQL, cosine similarity queries are fast enough, and schema management stays consistent with other tables

If you hit hundreds of millions of vectors and need specialized ANN tuning, switch then — not before


The Biggest Trap: AsyncSession Across Greenlets

LangGraph is async, and RAGAS evaluation also needs to run the retriever. But if your retriever uses a SQLAlchemy AsyncSession, there's a subtle problem waiting:

sqlalchemy.exc.MissingGreenlet: greenlet_spawn has not been called

Root cause: SQLAlchemy async sessions can't be passed across greenlets, but RAGAS spawns new greenlets when running evaluations

Fix: Use a sync psycopg2 connection in the retriever instead of AsyncSession

class ExpertScopedRetriever(BaseRetriever):
    """Uses sync psycopg2 — naturally compatible with LangGraph/RAGAS"""

    def _get_relevant_documents(self, query: str) -> list[Document]:
        embedding = self.embeddings.embed_query(query)
        # CTE + pgvector cosine similarity
        # Self-managed psycopg2 connection, not AsyncSession

The official documentation doesn't clearly describe this; it took a long time to debug


RAGAS Four Metrics

RAGAS breaks RAG quality into four quantifiable numbers:

Faithfulness

How much of the model's answer actually comes from the retrieved context?

Example: if RAG retrieved three passages and the model says something not in any of them, Faithfulness drops

This is the most important metric in education — you can't let AI invent answers

Answer Relevancy

Does the answer actually address the question? Models sometimes go off-topic; this metric catches that

Context Recall

How much of the ground truth information is covered by the retrieved context?

Low scores here usually point to chunking strategy or embedding quality problems — not necessarily the model

Context Precision

Of the retrieved passages, how many were actually useful? Is there a lot of irrelevant noise in the retrieval results?


Why Low Faithfulness Isn't Always the Model's Problem

This was the most important finding after building this

When people see Faithfulness at 0.6, they immediately start swapping models or tweaking prompts. But the problem might be somewhere else entirely

We designed a controlled experiment: the same test set, four combinations

With RAGWithout RAG
Strong modelVersion AVersion B
Weak modelVersion CVersion D

Compare A vs C: if the gap is small, the problem is in the RAG (chunk quality, retrieval strategy) Compare A vs B: if the gap is large, RAG is genuinely helping

This design lets you separate "RAG problems" from "model problems" so you don't keep fixing the wrong layer


Versioning Is What Makes Evaluation Meaningful

Once RAGAS produces numbers, you need to answer: "how much did this improve compared to the last version?"

That requires versioning the Expert's configuration

Every time you change a prompt, swap a chunk strategy, or adjust top_k — save it as a new version with a SHA256 hash

class ExpertVersion(Base):
    version = Column(Integer)
    config = Column(JSON)  # chunk_strategy, prompt_template, temperature, top_k...
    instruction_hash = Column(String)  # SHA256 of prompt

Without this, RAGAS scores can't be tied to specific configurations. Numbers improve but you don't know why


Early Observations

Evaluation data is still running, but a few initial findings:

  1. pgvector + 768-dim embeddings is enough: Search latency in the hundreds of milliseconds range, no issue for RAG use cases
  2. Chunk size matters a lot: Optimal chunk strategy for math problems vs reading passages is different — there's no universal value
  3. Faithfulness is very sensitive to prompt wording: Whether the instruction explicitly says "only answer based on the provided content" causes a large score difference

Summary

The biggest value of RAG + RAGAS is converting "does this AI answer well" — previously a gut feeling — into four trackable numbers

But the numbers don't tell you what to fix. Low Faithfulness → check the prompt. Low Context Recall → check the chunk strategy. Mapping the problem to the right layer is where this evaluation framework actually saves time

RAGRAGASLangChainpgvectorLangGraphEducation AI