A document conversion pipeline for RAG
Parse, chunk, embed, then JSONL is the whole RAG conversion pipeline; the language model sits after retrieval, not inside conversion.
Four stages, one artifact
Parse turns bytes into text and structure. Chunk cuts that text into overlapping passages sized for the embedder. Embed maps each passage to a vector. Export writes JSONL - one object per line - plus a sibling manifest. That pipeline is conversion. Everything after is storage and query.
Why the model is not a stage
Putting GPT or Claude inside the convert loop to 'understand' pages conflates generation with translation. You pay reasoning prices for work a sentence transformer finishes once, and you often lose a reusable corpus. Keep the frontier model on the answer path: retrieve five passages, then ask.
Where quality is decided
Chunk boundaries and parse fidelity set the ceiling. A mid-sentence cut or a flattened table cannot be fixed by a clever prompt later. OCR and speech are optional branches for scans and audio - they feed the same chunk stage, they do not replace it.
Idempotent by design
A pipeline you can re-run on the same file and get the same chunk ids and vectors is one you can trust in CI. Manifests that report per-file yields and skips make silent empty indexes visible. Without that, automation hides failure until users notice wrong answers.
Browser, API, agent - same pipeline
RAG Converter runs this pipeline in the tab for free on text-layer files, and on Fly via multipart POST for heavier work. MCP wraps the API for editors. The stages do not change; only where the bytes are allowed to go.