Batch convert documents for RAG
One folder in, one corpus out: batch conversion is how you build an index you can reload instead of reinventing per question.
The unit of work is the folder
Real knowledge is mixed: PDF, Word, slides, a spreadsheet of notes. Batch convert means that pile becomes one JSONL download (or one API response stream) with per-file counts in the manifest. A tool that only accepts one demo PDF is a demo.
Why batch beats per-question convert
Conversion is slow relative to retrieval. Doing it once per folder and querying many times is the bargain RAG sold. Re-sending pages to a model for every question reverses the bargain. Batch once; retrieve often.
What a good batch reports
Which files produced chunks, which were skipped and why - empty text layer, unsupported codec, size limit. Near-identical empty placeholders from silent failures poison hybrid search. Name the skips or the batch is lying.
Local batch first
Drop a folder in the browser converter when the material cannot leave the machine. You get the same JSONL shape as the API. Move to server batch for OCR-heavy scans, long audio or video when the trade is worth it.
From batch to AutoRAG
A successful manual batch is the template for automation: same paths, same engines, same output layout. Wire that command to a schedule or watcher and you have AutoRAG without changing the conversion math.