Incremental RAG indexing: only reconvert what changed
Re-embedding the whole library every night wastes money; incremental convert on new or changed files keeps cost and freshness aligned.
Full reindex is the expensive default
Teams that fear stale answers often rebuild everything on a cron. That works and it scales with the worst file in the tree every night. Incremental indexing asks a cheaper question first: did this file's content change?
What 'changed' means
Checksum or content hash beats mtime alone - copies and touch can fool clocks. When the hash differs, convert that file, upsert its chunks, delete old chunk ids for that source. When a file vanishes, remove its vectors. When nothing moved, sleep.
Chunk ids have to be stable
If every reconvert invents new ids, deletes and upserts become guesswork. Derive ids from source path plus chunk ordinal (or an equivalent stable scheme) so an incremental pass replaces the right rows. The manifest should make that mapping auditable.
OCR and media are special
A one-byte metadata edit should not re-run expensive speech or vision. Gate heavy engines on content hash of the media bytes, not on sidecar fields. Incremental without that gate is still a money fire.
Start simple
Even a 'convert files newer than last successful run' pass beats blind full rebuilds. Graduate to hashes when the library is large. RAG Converter's API and MCP convert tools fit either loop; the freshness policy is yours to own.