Recently I went deep on RAG in an education AI project and wired RAGAS evaluation directly into the pipeline
Here's what I learned, including some non-obvious parts
What RAG Is, and Why Education Needs It
RAG (Retrieval Augmented Generation) does one simple thing:
Before answering a question, search a knowledge base for relevant passages, then give those passages plus the question to the model
Why not just ask the model directly?
Models have a training cutoff and can't access private knowledge. In education this matters even more — textbooks, curricula, and worked solutions need version control. You can't have the model hallucinate answers
Stack Selection
The RAG stack for this project:
| Layer | Technology | Why |
|---|---|---|
| Vector DB | pgvector | Native PostgreSQL extension — no separate vector DB to operate |
| Embedding | Vertex AI text-embedding-004 | GCP ADC auth, no extra API key; 768 dimensions |
| Retriever | LangChain BaseRetriever | Native integration with LangGraph + RAGAS |
| LLM | LiteLLM wrapper | One-line model swap; supports Google, OpenAI, open source |
| Evaluation | RAGAS | Industry standard, four core metrics |
pgvector vs a standalone vector database
Many people's first instinct is Pinecone or Weaviate, but for this project there was no reason to operate an extra service. pgvector runs on top of PostgreSQL, cosine similarity queries are fast enough, and schema management stays consistent with other tables
If you hit hundreds of millions of vectors and need specialized ANN tuning, switch then — not before
The Biggest Trap: AsyncSession Across Greenlets
LangGraph is async, and RAGAS evaluation also needs to run the retriever. But if your retriever uses a SQLAlchemy AsyncSession, there's a subtle problem waiting:
sqlalchemy.exc.MissingGreenlet: greenlet_spawn has not been called
Root cause: SQLAlchemy async sessions can't be passed across greenlets, but RAGAS spawns new greenlets when running evaluations
Fix: Use a sync psycopg2 connection in the retriever instead of AsyncSession
class ExpertScopedRetriever(BaseRetriever):
"""Uses sync psycopg2 — naturally compatible with LangGraph/RAGAS"""
def _get_relevant_documents(self, query: str) -> list[Document]:
embedding = self.embeddings.embed_query(query)
# CTE + pgvector cosine similarity
# Self-managed psycopg2 connection, not AsyncSession
The official documentation doesn't clearly describe this; it took a long time to debug
RAGAS Four Metrics
RAGAS breaks RAG quality into four quantifiable numbers:
Faithfulness
How much of the model's answer actually comes from the retrieved context?
Example: if RAG retrieved three passages and the model says something not in any of them, Faithfulness drops
This is the most important metric in education — you can't let AI invent answers
Answer Relevancy
Does the answer actually address the question? Models sometimes go off-topic; this metric catches that
Context Recall
How much of the ground truth information is covered by the retrieved context?
Low scores here usually point to chunking strategy or embedding quality problems — not necessarily the model
Context Precision
Of the retrieved passages, how many were actually useful? Is there a lot of irrelevant noise in the retrieval results?
Why Low Faithfulness Isn't Always the Model's Problem
This was the most important finding after building this
When people see Faithfulness at 0.6, they immediately start swapping models or tweaking prompts. But the problem might be somewhere else entirely
We designed a controlled experiment: the same test set, four combinations
| With RAG | Without RAG | |
|---|---|---|
| Strong model | Version A | Version B |
| Weak model | Version C | Version D |
Compare A vs C: if the gap is small, the problem is in the RAG (chunk quality, retrieval strategy) Compare A vs B: if the gap is large, RAG is genuinely helping
This design lets you separate "RAG problems" from "model problems" so you don't keep fixing the wrong layer
Versioning Is What Makes Evaluation Meaningful
Once RAGAS produces numbers, you need to answer: "how much did this improve compared to the last version?"
That requires versioning the Expert's configuration
Every time you change a prompt, swap a chunk strategy, or adjust top_k — save it as a new version with a SHA256 hash
class ExpertVersion(Base):
version = Column(Integer)
config = Column(JSON) # chunk_strategy, prompt_template, temperature, top_k...
instruction_hash = Column(String) # SHA256 of prompt
Without this, RAGAS scores can't be tied to specific configurations. Numbers improve but you don't know why
Early Observations
Evaluation data is still running, but a few initial findings:
- pgvector + 768-dim embeddings is enough: Search latency in the hundreds of milliseconds range, no issue for RAG use cases
- Chunk size matters a lot: Optimal chunk strategy for math problems vs reading passages is different — there's no universal value
- Faithfulness is very sensitive to prompt wording: Whether the instruction explicitly says "only answer based on the provided content" causes a large score difference
Summary
The biggest value of RAG + RAGAS is converting "does this AI answer well" — previously a gut feeling — into four trackable numbers
But the numbers don't tell you what to fix. Low Faithfulness → check the prompt. Low Context Recall → check the chunk strategy. Mapping the problem to the right layer is where this evaluation framework actually saves time
