This is Part 1 of a 5-part series documenting my Master of Science in Software Engineering Capstone project at Quantic: a deployed, agentic RAG system that answers natural-language questions about European electricity markets. Follow along as I go from problem statement to production deployment.

A capstone project is a culminating academic assignment that typically takes place at the end of a degree program. So, this is sort of my master’s thesis.


Why power markets need a better way to ask questions

If you’ve ever worked a trading desk, you know the drill. Something moves in the market; a price spike, a sudden change in the balance between supply and demand, and the first question is always the same: why?

The answer is usually scattered across half a dozen places. A structured outage report from a transmission system operator. A press release buried in a regulatory filing feed. A news article that mentions a plant name once, in passing, three paragraphs in. None of these sources talk to each other. None of them are built to answer a plain question like “why did DE-LU prices spike on Tuesday?” You have to go hunting, cross-reference by hand, and hope you didn’t miss the one document that actually explains it.

I’ve spent years on the other side of that problem. Through independent work building a Nordic power dashboard and a Monte Carlo framework for PPA valuation, I got intimately familiar with how much manual effort goes into synthesizing market-moving information, and how much of that effort is genuinely repetitive. It’s not that the information is unavailable. It’s that no one has built the tool that reads it all for you and gives you a straight, cited answer.

That gap is the starting point for this capstone.

The vision: a research analyst that never sleeps and always cites its sources

The project is called the Agentic Power-Market Analyst: a retrieval-augmented generation (RAG) system that takes natural-language questions about European electricity markets and answers them grounded in a live, versioned corpus of real sources: structured outage data, transmission system operator UMM notices, and curated energy news.

A few design principles matter a lot here, and they’re what make this project genuinely fun to build:

  • Every answer is cited. This is a market-analysis tool, not a content generator. If the system tells you why a bidding zone saw a supply squeeze, it points to the exact outage notice or news article behind that claim.
  • It’s agentic, not just retrieval. Beyond searching a document corpus, the system is being built to reach out to live data sources; like ENTSO-E’s structured market data, mid-conversation, the way a human analyst would pull up a live dashboard instead of relying only on what’s already been written up.
  • It’s deployed. It’s a live, hosted system with a real URL, built to be used.

Starting narrow, on purpose

The initial scope covers three bidding zones: DE-LU (Germany-Luxembourg), DK1 (Western Denmark), and NO2 (Southern Norway), chosen because they’re liquid, well-documented, and structurally distinct enough to stress-test the system’s design. The intent from day one has been to expand outward: eventually covering all of Europe, or letting users select the zones that matter to them.

That “start narrow, prove it works, then generalize” approach isn’t just good engineering hygiene, it’s also the only way to keep a capstone project shippable on a semester timeline. A pan-European system with no ingestion boundary and no clear evaluation set is a research agenda, not a deliverable. Three zones with a clean expansion path is a product.

What’s ahead in this series

This series will trace the project from vision to working system, including the parts that didn’t go as planned, because those are usually the most interesting parts:

  • Part 2: The architecture pivot: why I abandoned ingestion-time geography filtering in favor of pan-European ingestion with query-time zone filtering, and what that taught me about over-engineering early.
  • Part 3: Building the corpus: wrangling structured outage data and news sources into something a RAG pipeline can actually trust.
  • Part 4: Making it agentic: adding a live ENTSO-E data tool so the system can go beyond its static corpus.
  • Part 5: Deployment, evaluation, and the lessons that only show up once something is live.

I’m building a real, cited, deployed AI system in a domain I know from the trading side, and I’m documenting the engineering tradeoffs.

Next up: the architecture pivot that reshaped the whole project.

Leave a Reply

Trending

Discover more from Convergence Point

Subscribe now to keep reading and get access to the full archive.

Continue reading