RCRAG Converter

File-to-embeddings automation

Turning files into embeddings is a repeatable job, not a model session - automate the convert, reserve the frontier model for the answer.

Embeddings are a product of convert

A file does not become searchable by chatting with it. It becomes searchable when passages exist as vectors. File-to-embeddings automation is scheduled or event-driven conversion that writes those vectors (and the text that backs them) somewhere queryable.

Small models for the map, large for the reply

Sentence transformers map meaning into a few hundred dimensions cheaply. Frontier models generate answers. Automating file-to-embeddings with a frontier model is the cost and privacy mistake: you ship the whole file to a vendor to do work a local MiniLM-class model can finish.

What to store beside the vector

Id, text, source path, page or offset when available, and how the text was obtained - parsed, OCR, ASR, or vision description. Automation that drops metadata makes citations impossible and debugging a coin flip.

Throughput is an ops problem

Queue files, bound concurrency, retry transient failures, dead-letter poison files. The embedder is rarely the hard part once chunking is sane; the hard part is not losing a Tuesday's Dropbox sync on a single corrupt PDF.

Where RAG Converter fits

Browser path: folder to JSONL with optional on-device embeddings, nothing uploaded. API path: multipart convert for scale and media engines. MCP: same API for agent-driven automation. All three emit the file-to-embeddings step as a corpus, not as a chat transcript.

More on converting for RAG