Your firm's IT or compliance reviewer asks a fair question before any AI tool touches client files: which outside companies will actually see this data? The honest answer for almost every AI product is that more than one company is involved, even if only one name is on the invoice. A sub-processor is any outside vendor that the AI company sends your data to in order to make the product work — an embeddings provider, a model provider, sometimes a hosting company. Reading a sub-processor list well means checking three things for each name on it: what data it receives, why that data is necessary, and what it promises not to do with it.
This matters more for legal work than most software categories, because privileged and confidential material carries duties that follow the document, not just the vendor relationship. A firm that hands work to an unnamed vendor has effectively delegated part of its confidentiality obligation without knowing to whom.
What is a sub-processor in an AI vendor's stack?
A sub-processor is a company that processes your data on behalf of the vendor you actually signed up with, without you having a direct contract with that sub-processor. In a legal AI tool built on retrieval-augmented generation, this almost always means at least two: a provider that turns your document text into searchable embeddings, and a provider whose language model drafts the answer from the passages retrieved. Neither is optional, which is exactly why the list exists: not to alarm you, but to tell you precisely who else is in the chain.
A vendor that avoids naming its sub-processors is not avoiding risk, it is avoiding your ability to evaluate the risk. The absence of a named list is itself information.
The alternative is not risk-free either: refusing any tool with sub-processors usually just pushes the same behavior underground, with a paralegal pasting passages into a general chatbot that has no sub-processor list, no encryption commitment, and no audit trail at all. The real question is not whether a sub-processor exists, but whether the vendor names it and limits what it sees.
Which outside vendors touch your documents when you use Vorticel?
Vorticel uses exactly two sub-processors, both named on its security page: Voyage AI (a MongoDB company) for embeddings, and Anthropic for drafting answers. Each receives only the minimum data its step requires, and the original file itself never leaves Vorticel's encrypted storage. What moves is extracted text and, later, a small number of retrieved passages, described in more detail in where your uploaded documents actually go.
| Step | What is sent | Sent to | Why it is necessary |
|---|---|---|---|
| Upload | Nothing external yet | Vorticel's own encrypted storage | Document is tagged to your firm and encrypted before anything else happens |
| Embedding | Extracted document text, split into passages | Voyage AI | Converts text into the numeric form that makes search possible |
| Answering | Only the small number of passages retrieved for one question | Anthropic | Drafts the answer from retrieved passages only, never the full document |
| Training | Nothing, by policy | No one | Your documents, questions, and answers are never used to train a model |
Notice what does not appear in that table: a full document ever reaching either sub-processor, or either one retaining data beyond the request that needed it. That is the pattern a sub-processor list should let you verify, not just assert.
This is also why the retrieval step matters as much as encryption at rest: even a vendor with strong encryption still sends something to a sub-processor at query time. The only real controls left are how much is sent, to how many parties, and whether either one keeps a copy longer than the single request needs.
What should a compliance reviewer check on a vendor's sub-processor list?
Check five things, in this order, before approving any AI vendor for privileged material. Each one is a question the vendor should be able to answer without hedging.
- Get the full list of named sub-processors, the actual company names, not a category description like 'cloud infrastructure partners.'
- For each one, confirm exactly what data it receives: the whole document, extracted text, or only retrieved passages.
- Confirm whether any sub-processor uses your data to train its own models, and get that in writing, not implied by the primary vendor's own policy.
- Ask how a sub-processor is added in the future. You want advance notice, not a silent update to a terms page.
- Cross-check the list against the vendor's own description of its data flow. A mismatch between the two is a bigger red flag than either document alone.
This is the same ground covered from the buyer's side in questions to ask an AI vendor before uploading client documents, a sub-processor list is where several of those questions get answered on paper instead of in a sales call.
Why does a no-training clause matter if sub-processors still see your text?
Because a no-training promise is about use, not about access. A sub-processor can receive your document text for the narrow purpose of generating an embedding or drafting an answer, and still never use that text to improve its own models. Those are two separate commitments, and a vendor's sub-processor list should make both explicit rather than letting one imply the other. When you evaluate the clause, look for whether it names the sub-processor specifically, stating that the embeddings provider itself does not train on submitted text, rather than only describing the primary vendor's own policy.
The cleanest way to get an unambiguous answer is to ask the vendor for the sentence in writing: name the sub-processor and state, without qualification, that it does not train on data submitted through the product. A vendor that can produce that exact sentence has clearly asked its own sub-processors the same question you are asking it.
If your firm handles work under an ethical duty of confidentiality, treat every product data flow as if it will eventually be described in an engagement letter or a client audit. The tools that hold up under that description are the ones whose sub-processor list already matches what actually happens in production, not just what the marketing page says. Vorticel's own documentation describes this same data flow at the field level, for teams that want to check it against their own client-confidentiality obligations before approving the tool.
The most common objection at this stage is reasonable: your firm already vets vendors itself, so why does this vendor's list matter more than your own review? It does not replace your review, it is the input your review needs. A firm's own vendor-diligence process is only as good as the disclosures it has to work with, and a vendor that names its sub-processors and states the minimum data each one receives has given your reviewer something concrete to sign off on, instead of a policy paragraph to take on faith.
How do you start checking a vendor's sub-processor list today?
Ask for the list in writing before you upload a single document, not after. If you are evaluating Vorticel specifically, its full data flow and both named sub-processors are on the security page today, in the same detail shown above. Start a free trial and read that page side by side with your firm's own document-handling policy. It takes a few minutes and tells you, before any client file is involved, exactly who would see what.