RCRAG Converter

Batch convert documents for RAG

One folder in, one corpus out: batch conversion is how you build an index you can reload instead of reinventing per question.

The unit of work is the folder

Real knowledge is mixed: PDF, Word, slides, a spreadsheet of notes. Batch convert means that pile becomes one JSONL download (or one API response stream) with per-file counts in the manifest. A tool that only accepts one demo PDF is a demo.

Why batch beats per-question convert

Conversion is slow relative to retrieval. Doing it once per folder and querying many times is the bargain RAG sold. Re-sending pages to a model for every question reverses the bargain. Batch once; retrieve often.

What a good batch reports

Which files produced chunks, which were skipped and why - empty text layer, unsupported codec, size limit. Near-identical empty placeholders from silent failures poison hybrid search. Name the skips or the batch is lying.

Local batch first

Drop a folder in the browser converter when the material cannot leave the machine. You get the same JSONL shape as the API. Move to server batch for OCR-heavy scans, long audio or video when the trade is worth it.

From batch to AutoRAG

A successful manual batch is the template for automation: same paths, same engines, same output layout. Wire that command to a schedule or watcher and you have AutoRAG without changing the conversion math.

More on converting for RAG