RCRAG Converter

Automated API PDF parsing for Salesforce CRM intelligence

Automated API PDF parsing for Salesforce CRM intelligence converts each PDF attached to a record into searchable passages, so the assistant answers from the file rather than from fields someone typed.

How does automated API PDF parsing feed Salesforce CRM intelligence?

Automated API PDF parsing for Salesforce CRM intelligence converts each attached PDF into passages when the file is saved, so the assistant can cite the file instead of guessing from record fields.

Where the PDF sits in Salesforce

Salesforce keeps the file on the lead, account, or opportunity as a ContentVersion or a classic Attachment. The typed fields on that record are already searchable. The PDF is not, until automated API PDF parsing copies its words into passages keyed to the same record. CRM intelligence that only reads the fields is intelligence about what a person typed, not about what the file says.

What a PDF actually yields

A PDF with a text layer stores words as text, and that layer becomes passages. A scanned PDF stores a photograph of the page, so the same extension yields nothing until OCR runs. Two-column contracts and footnotes can extract out of reading order, which is why a sample of the JSONL should be read before a whole library is indexed.

The API call that feeds Salesforce

An Apex trigger, a record-triggered Flow, or a Platform Event can POST those bytes when the file is saved. The response is JSONL, one passage per line, with the record id in metadata for the retrieval store beside Salesforce. Asking a model to reread the PDF inside chat for every question never builds that index. Automated API PDF parsing runs when the file is saved, once, and later questions retrieve.

What the answer is allowed to cite

After a clean parse, an answer should cite the passage and the PDF it came from. A file that yields no text is a skip, named as a skip, not a blank passage that crowds real hits. The parsers are Remade with Rust crates, and RAG Converter is part of MATA. The longer note on this format is RAG for PDF.

More on converting for RAG