Keeping a RAG Index Fresh
Keep retrieval in sync with changing documents: updates, deletions, versions, permissions and re-embedding.
A taste of a lesson
We updated a policy, but the bot sometimes still quotes the old wording. What went wrong?
Most likely the old chunks were never removed. If updates insert new chunks without deleting the previous ones, both versions live in the index and either can be retrieved. Key every chunk by document id, and on update delete all chunks for that id before inserting the new set, ideally in one operation. Also check for an answer cache holding the old reply. Then search for a phrase that exists only in the old wording. Does anything still come back?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Choose between full rebuilds, incremental updates and reconciliation
- Detect changes with timestamps, hashes and change feeds
- Design stable chunk ids so updates replace old content cleanly
- Handle deletions, permission revocations and document versions reliably
- Migrate to a new embedding model without downtime and monitor freshness
Lesson plan
- 1 How documents change List change types and set freshness targets per source. Start
- 2 Rebuild or increment Choose a sync strategy that fits size and change rate. Start
- 3 Change detection and stable ids Detect changes reliably and replace chunks cleanly. Start
- 4 Deletions, permissions and versions Give the critical change types the care they need. Start
- 5 Re-embedding migrations Switch embedding models or chunking without downtime or mixing. Start
- 6 Monitoring freshness Notice staleness before users do. Start
Try asking
About this tutor
For experienced developers running retrieval over documents that change: policies revised monthly, wikis edited daily, files deleted, permissions changed. You learn to choose between full rebuilds and incremental updates, detect changes with timestamps, hashes and change feeds, design stable chunk ids so updates replace rather than duplicate, and treat deletions and permission revocations as the critical path. You also plan a re-embedding migration when switching embedding models, handle multiple versions of the same document, and monitor freshness so stale answers are noticed before users complain.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Takes retrieval systems from demo to dependable: parsing, citations, freshness, retrieval evaluation and debugging
9 tutors 267 lessons taught Sample
Most RAG demos work on the ten documents someone picked. I teach what happens after that: scanned PDFs, tables, documents that change every week, answers that cite the wrong page and users who ask things the documents never covered. My background is in document processing and internal knowledge tools, so I am practical about formats and sceptical of any setup...
See Fumiko's profile and tutorsMore like this
Other tutors on the same or nearby topics.