RAG, short for retrieval-augmented generation, makes an AI model answer from your documents instead of from memory. Before the model answers, the system finds the passages most relevant to the question and hands them over, with an instruction to answer from those passages only. That cuts down on made-up answers, and it lets the model use documents it never saw in training, such as a company's own files or a law that changed last month.
A RAG demo is quick to build. A system you can rely on takes more work, because you have to show that it finds the right passages and answers correctly from them.
I built one over the EU AI Act. You ask a question, and it answers with the exact article, recital or annex behind each statement, or says that the Act does not answer the question. The code and every evaluation run are on GitHub.
Step 1: Scope it
A system meant for every user, every question and every document tends to be built badly or never finished, so I started narrow:
- Users: people who need to know what the AI Act requires of their company, such as compliance teams.
- Questions: what the Act itself says: definitions, obligations, deadlines, and who they apply to. Not legal advice, and not how a particular country applies the Act.
- Documents: one. The regulation as published in the EU's Official Journal in July 2024, with 180 recitals that explain its purpose, 113 articles that set the rules, and 13 annexes.
Two requirements follow from who the users are. Every statement must cite the provision it comes from, so a reader can check it. And when the Act does not answer a question, the system must say so instead of guessing.
Step 2: Write the test first
The test is a set of questions with known answers: for each question, the part of the law that answers it. I wrote it before tuning anything, because it is how I know whether a change helped.
It has 54 questions. 44 are answered by the Act, each labelled with the provisions that answer it: "What does 'AI literacy' mean?" is answered by Article 3(56). The other 10 are questions the Act does not answer, such as which authority supervises it in Norway, because a system that never has to refuse cannot show that it knows when to. I tuned on 37 of the questions and held 17 back for the final test.
One rule matters more than the others: the answers in the test come from the source documents, never from the system being tested. A test built from the system's own output can only agree with it, mistakes included.
Step 3: Build the retrieval and measure it
Steps 3 and 4 build the pipeline below. It is not an agent: code decides every step, and the language model is called once, to write the answer.
Retrieval finds the passages. The law is cut into one piece per provision, so an answer can cite exactly one of them. The code that does the cutting also sets the label that every later step relies on, so it is the first thing I check: problems in the documents cause bad answers that no amount of tuning fixes. For each question, the system searches by meaning and by keywords, and a second model, a reranker, puts the best candidates in order.
I measure retrieval by how often the piece that answers the question comes first (Recall@1), and how often it is in the top three (Recall@3), which is what the answer step sees. The test matches pieces to provisions by the text they contain, not by their labels, so a mislabelled piece cannot count as a hit.
In the final system, the right piece comes first for 85% of the test questions, and it is in the top three for all of them.
Step 4: Build the answers and check them
The answer step gives Claude the top three pieces and one instruction: answer from these passages only, cite the provision behind every statement, and if the passages do not answer the question, say so.
Every answer is then checked:
- Citations: every cited provision is checked against the pieces the model was given, as a legal reference rather than as text.
- Declines: questions the Act cannot answer must be declined, and questions it can answer must not be.
- Answer quality: where code cannot judge, such as whether an answer is complete, a second model grades it against the Act's own text and against the passages the model was given.
A small viewer shows each answer next to its question, the retrieved pieces and the citation check, which is where failures are easiest to spot.
In the final system, no citation points to the wrong provision, and every question the Act cannot answer is declined.
Step 5: Change one thing at a time
With the test and the checks in place, every change is an experiment: change one thing, run the whole test again, and keep the change only if the numbers improve.
The change that mattered most was about recitals, the parts of the Act that explain its purpose. With every provision in its own piece, recitals crowded the top of the results, often above the article that actually answers the question. Most likely because recitals read like the questions people ask, while articles are written in legal drafting.
- One piece per provision: the right piece came first 69% of the time on the test questions and 52% on the development questions.
- A wider pool of candidates for the reranker did not help on its own (61% and 48%), but the next change needed it: the reranker can only move an article up if the article is among the candidates.
- Ranking recitals below articles raised it to 85% and 68%. I tuned it on the development questions only, then ran it once on the test questions.
The second run looked like a step backwards, yet the third could not work without it. Measured together, that would have been invisible.
The result
The final system, on the 17 test questions and the 37 development questions:
| Test | Development | |
|---|---|---|
| Right piece first (Recall@1) | 85% | 68% |
| Right piece in the top three (Recall@3) | 100% | 81% |
| Citations pointing at the wrong provision | 0 | 0 |
| Completeness, judged against the Act | 0.96 | 0.81 |
| Faithfulness to the retrieved passages | 1.00 | 0.98 |
| Unanswerable questions declined | 4 of 4 | 6 of 6 |
| Answerable questions declined | 0 of 13 | 2 of 31 |
| Cost per question | 1.2 cents | 1.2 cents |
The two declined development questions are one retrieval miss, and one borderline question on whether Article 4 sets a number of training hours, where the system says the Act does not answer instead of saying "no". With 13 answerable test questions, one question moves a test score by 8 points, so small differences are noise.
What the checks caught
The checks in steps 2 and 3 are there because each caught a real problem while I was building this system.
The labels. Checking each piece's label against the text inside it showed that 39% of the text sat under a label naming a different provision, and another 16% had no label at all:
References to other articles that happened to start a line in the PDF were read as the start of a new article, so recitals 62 to 180 became one piece labelled "Article 39". And Article 3's numbered definitions were read as recitals, so the definition of AI literacy, Article 3(56), was labelled "Recital 56".
A test that agreed with the system. An answer key built with the system's help, an LLM picking the right passages from the system's own results, took those labels as they were. For a question on AI literacy, it listed the passage labelled "Recital 56" as a right answer. So a citation to "Recital 56" counted as correct, and that test reported perfect retrieval.
Limits
- The corpus is the 2024 text. The Digital Omnibus on AI (Regulation (EU) 2026/1744), in force since 27 July 2026, changes some application dates and is not included. The app says so next to every answer.
- English only. Another language needs a multilingual reranker and a new test.
- The judge is a model, and can be wrong. The citation, decline and retrieval numbers do not depend on it.
- The test's answers have not yet been reviewed by a legal expert.
The code and every evaluation run are on GitHub: eu-ai-act-rag. A companion post walks through the research agent I built on the same regulation.