Debugging a RAG App That Answers Badly
Trace each bad RAG answer to its real cause, from missing documents to ignored passages, and fix the most common.
A taste of a lesson
Our bot answers the first question well, but 'and what about part-time staff?' gets nonsense. Why?
Look at what was actually searched for. A follow up like that, searched on its own, has almost no meaning: no mention of the topic from the first question, such as holiday entitlement. Retrieval then returns random part time content. The usual fix is to rewrite the follow up into a standalone query using the conversation history before searching, for example 'holiday entitlement for part time staff'. Then log the rewritten query to check it. What query does your trace show for that turn?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Reproduce a bad answer from a complete trace
- Walk a fixed cause list from content gaps to citation errors
- Tell retrieval failures from generation failures with quick checks
- Keep a categorised failure log and fix the largest category first
- Turn fixed failures into permanent evaluation cases
Lesson plan
- 1 Reproduce from a full trace Capture everything needed to replay a bad answer. Start
- 2 Content and parsing problems Rule out missing documents and broken extraction first. Start
- 3 Retrieval misses Find out why the right chunk was not retrieved. Start
- 4 Ranking and generation problems Diagnose good evidence that was cut off, ignored or misread. Start
- 5 Failure logs and priorities Categorise failures and fix the most common cause first. Start
- 6 Keeping fixes fixed Protect improvements with evaluation cases. Start
Try asking
About this tutor
For developers whose retrieval augmented app gives wrong, vague or unsupported answers and who are tired of tweaking prompts at random. You learn a systematic method: reproduce the failure from a full trace, then walk down a fixed list of causes, from the answer not existing in the corpus, through parsing, chunking, filters and ranking, to the model ignoring or misreading good evidence. You practise on real failure cases, keep a failure log with categories and counts, fix the most frequent cause first, and turn every fixed case into a permanent evaluation example so it stays fixed.
Reviews
4.7
3 ratingsSample
- Darius N.Sample
The trick of rerunning with only the correct passage told us immediately whether to fix retrieval or generation. Practical lessons.
- Lucia M.Sample
The cause list in order is now pinned in our team channel. Half our bad answers were content gaps, not model problems. We stopped rewriting the prompt every day.
- Sophie R.Sample
Follow up question rewriting fixed our worst complaints. The failure log habit keeps us honest about priorities.
About the teacher
Takes retrieval systems from demo to dependable: parsing, citations, freshness, retrieval evaluation and debugging
9 tutors 267 lessons taught Sample
Most RAG demos work on the ten documents someone picked. I teach what happens after that: scanned PDFs, tables, documents that change every week, answers that cite the wrong page and users who ask things the documents never covered. My background is in document processing and internal knowledge tools, so I am practical about formats and sceptical of any setup...
See Fumiko's profile and tutorsMore like this
Other tutors on the same or nearby topics.