This is the final part of a 5-part series documenting my MSSE Capstone project at Quantic: a deployed, agentic RAG system that answers natural-language questions about European electricity markets. Part 4 covered the agent layer and its bugs — this one covers getting it live, proving it works, and the last complication that showed up right at the end.
Getting it out of my dev environment
A system that only runs on my laptop isn’t a deliverable, it’s a demo. The app is built on FastAPI, backed by Neon Postgres with pgvector for the embedding store, deployed on Render’s free tier, with GitHub Actions handling CI. None of that is exotic, and that was deliberate — the interesting engineering in this project is in the retrieval design and the agent loop, not in the deployment plumbing. I wanted infrastructure that would stay out of the way, not become its own project.
Two terminals became the standard working setup by the end: one running uvicorn app:app --reload for the live server, one free for scripts, diagnostics, and eval runs. It’s a small thing, but it’s the kind of small thing that saves real time across a few hundred iterations.
Proving it works: the eval harness
Throughout this series I’ve mentioned the eval harness catching bugs — the zone-blind guardrail, the relevance-gate regressions, the citation-bracket bug. It’s worth saying plainly what that harness actually is: run_eval.py, run against a versioned, 37-question gold set, scoring five things — whether the system correctly refuses when it should, whether citations actually match the claims they support, whether facts are accurate, whether the right tool gets called, and whether the answer is generally plausible.
That gold set grew over the project, and grew carefully — every new question got added because it exposed a real gap, not to pad a number. It’s also the reason I trust the final numbers: they weren’t tuned to a small hand-picked sample, they’re the result of a set that kept catching new failure modes as the system got more capable.
One concrete example of the harness earning its keep: a retrieval-k ablation sweep across k=4, 6, 8, and 10 against the full gold set, to settle empirically — rather than by intuition — how many chunks retrieval should pull back. k=6 came out on top. It’s a small parameter, but it’s the difference between “I think 6 feels right” and “I ran the numbers and 6 is right.”
The last complication: adding a fourth zone
The system launched scoped to three zones — DE-LU, DK1, NO2 — for the reasons covered in Part 2. Late in the project, I added a fourth: SE3, Sweden, as a stretch goal. Because zone had been designed as a query-time filter rather than an ingestion-time boundary, back in Part 2, adding SE3 to ingestion and the config was genuinely straightforward.
What wasn’t straightforward was finding every place in the codebase that had quietly assumed there would only ever be three zones. Two bugs surfaced specifically because of this addition:
- A static refusal string, written back when three zones was the whole universe, that simply didn’t know a fourth zone existed.
- A guardrail prompt that listed the valid zones by name instead of reading them from
config.ZONES— meaning the single source of truth for “what zones exist” and the text actually shown to the guardrail had quietly drifted apart.
Both are the same category of bug: correct when written, wrong the moment the assumption they were built on changed. Neither would have been caught without deliberately testing the fourth zone end-to-end rather than assuming the config change was sufficient on its own. It’s a good closing example for the whole project, honestly — the system’s architecture handled the hard part of expansion (query-time zone filtering) exactly as designed, and the bugs that remained were in the parts that hadn’t been forced to think generically yet.
The final numbers
After all of the above — the pivots, the corpus work, the agent debugging, the fourth zone — here’s where the system landed on the 37-question gold set:
| Metric | Score |
|---|---|
| Refusal correctness | 0.97 |
| Citation hit rate | 0.86 |
| Fact match | 0.76 |
| Tool selection | 1.00 |
| Plausibility | 1.00 |
Tool selection and plausibility at a clean 1.00 say the agent reliably knows when to reach for a live number versus the corpus — which was the whole point of Sprint 3. Refusal correctness at 0.97 says the guardrail work in Part 3 and the zone-blindness fix in Part 4 actually paid off. Fact match at 0.76 is the most honest number on the board — the hardest thing to get exactly right, and the one with the most room left to improve if this project continued.
Where it ended up
The system is live, deployed, and covers four bidding zones with cited, grounded answers to natural-language questions about European power markets — built from a semester of picking apart a problem I’d lived on the other side of as a trader, pivoting twice before the architecture was right, and chasing down bugs that ranged from a stray retry loop to a config file that hadn’t caught up with its own zone list.
If there’s one thread running through all five parts, it’s this: almost nothing that shaped the final system was obvious on day one. The trading-signal idea looked right until it didn’t. The ingestion-time zone boundary looked clean until a cross-border UMM broke it. The guardrail looked correct until an eval run caught it quietly failing. The pattern that saved the project every time wasn’t getting things right the first time — it was building a harness rigorous enough to notice when they weren’t, and being willing to go find out why.
That’s the project. Thanks for following along.





Leave a Reply