RAG conversion is ETL for unstructured files
RAG conversion is ETL for unstructured files - extract, transform (chunk and embed), load JSONL - then query like any other warehouse table.
Unstructured still needs a pipeline
Data teams already run ETL for databases and events. PDFs and decks get a special exemption and land in a chatbot instead. That exemption is why RAG projects stall: there is no extract-transform-load, only hope. Conversion is the ETL. Retrieval is the query layer.
Extract
Pull text and structure from Office, text-layer PDF, HTML, CSV. Branch to OCR or speech when the bytes are pixels or audio. Record how each passage was obtained. Extraction that invents prose you cannot find in the file is not ETL - it is hallucination upstream.
Transform
Chunk on paragraph and sentence boundaries with overlap sized for the embedding model. Embed with a model you can re-run. Reject or quarantine empty yields. Transform is where chunking strategy and OCR quality set the ceiling for every later SQL-like vector lookup.
Load
JSONL into your vector store or lake. Idempotent loads keyed by source and chunk id. Manifests as run reports. The same discipline you use for warehouse loads applies: no silent truncates, no mystery row counts.
Then query - do not re-ETL per question
Once loaded, questions are retrieval plus generation. Re-sending the PDF to a frontier model for each ask is re-running extract on every SELECT. Automate ETL on file change; spend model budget on the answer. Browser or API convert is the extract-transform step; your store is the load.