RCRAG Converter

Automate file management for RAG

Folder watchers and batch convert beat drop-files-in-a-chatbot when the knowledge base is a living directory, not a one-shot upload.

The knowledge is already in folders

Most organisations already manage files: shared drives, Dropbox, S3, a wiki export cron. RAG does not need a second filing system. It needs those trees to become passages on a schedule. Automating file management for RAG means treating the directory as the source of truth and conversion as the sync step - the same idea as syncing photos, applied to embeddings.

Chat upload is not file management

A chatbot that accepts a PDF stores a session copy and answers from it. Useful for one document. It does not version with your share drive, does not tell you which of forty policies failed to parse, and does not emit a corpus you can load into your own vector database. File management for RAG starts with the folder you already trust.

What to automate first

Stable paths: contracts/, policies/, runbooks/. Ignore temp and inbox noise. Convert text-layer documents first; park scans for OCR. Emit JSONL per batch with a manifest that names skips. Then wire delete/rename so removed files leave the index. Skipping that last step leaves ghosts that outrank real answers.

Human jobs that stay human

Choosing what is in scope, classifying confidential paths that must stay local, and reviewing OCR on critical scans. Automation does the boring convert. A person still owns the allowlist. Mixing those jobs is how a watcher cheerfully indexes something that was never meant to leave the machine.

A practical split

Sensitive folders: convert in the browser or on a locked host, no third party. High-volume or media-heavy trees: API convert with prepaid credit. Agents can call the same convert via MCP once a key is in env. The file tree stays yours; only the conversion step moves to the tool that fits.

More on converting for RAG