How to Read a Legal AI Vendor's Sub-Processor List

In shortA sub-processor is any outside company an AI vendor sends your data to in order to deliver the product — an embeddings provider, a model provider, a hosting company. Reading the list means checking what each one receives, why, and whether it trains on your data. A vendor that names its sub-processors and states the minimum data each one gets has already answered the hardest compliance question.

Your firm's IT or compliance reviewer asks a fair question before any AI tool touches client files: which outside companies will actually see this data? The honest answer for almost every AI product is that more than one company is involved, even if only one name is on the invoice. A sub-processor is any outside vendor that the AI company sends your data to in order to make the product work — an embeddings provider, a model provider, sometimes a hosting company. Reading a sub-processor list well means checking three things for each name on it: what data it receives, why that data is necessary, and what it promises not to do with it.

This matters more for legal work than most software categories, because privileged and confidential material carries duties that follow the document, not just the vendor relationship. A firm that hands work to an unnamed vendor has effectively delegated part of its confidentiality obligation without knowing to whom.

What is a sub-processor in an AI vendor's stack?

A sub-processor is a company that processes your data on behalf of the vendor you actually signed up with, without you having a direct contract with that sub-processor. In a legal AI tool built on retrieval-augmented generation, this almost always means at least two: a provider that turns your document text into searchable embeddings, and a provider whose language model drafts the answer from the passages retrieved. Neither is optional, which is exactly why the list exists: not to alarm you, but to tell you precisely who else is in the chain.

A vendor that avoids naming its sub-processors is not avoiding risk, it is avoiding your ability to evaluate the risk. The absence of a named list is itself information.

The alternative is not risk-free either: refusing any tool with sub-processors usually just pushes the same behavior underground, with a paralegal pasting passages into a general chatbot that has no sub-processor list, no encryption commitment, and no audit trail at all. The real question is not whether a sub-processor exists, but whether the vendor names it and limits what it sees.

Which outside vendors touch your documents when you use Vorticel?

Vorticel uses exactly two sub-processors, both named on its security page: Voyage AI (a MongoDB company) for embeddings, and Anthropic for drafting answers. Each receives only the minimum data its step requires, and the original file itself never leaves Vorticel's encrypted storage. What moves is extracted text and, later, a small number of retrieved passages, described in more detail in where your uploaded documents actually go.

StepWhat is sentSent toWhy it is necessary
UploadNothing external yetVorticel's own encrypted storageDocument is tagged to your firm and encrypted before anything else happens
EmbeddingExtracted document text, split into passagesVoyage AIConverts text into the numeric form that makes search possible
AnsweringOnly the small number of passages retrieved for one questionAnthropicDrafts the answer from retrieved passages only, never the full document
TrainingNothing, by policyNo oneYour documents, questions, and answers are never used to train a model

Notice what does not appear in that table: a full document ever reaching either sub-processor, or either one retaining data beyond the request that needed it. That is the pattern a sub-processor list should let you verify, not just assert.

This is also why the retrieval step matters as much as encryption at rest: even a vendor with strong encryption still sends something to a sub-processor at query time. The only real controls left are how much is sent, to how many parties, and whether either one keeps a copy longer than the single request needs.

What should a compliance reviewer check on a vendor's sub-processor list?

Check five things, in this order, before approving any AI vendor for privileged material. Each one is a question the vendor should be able to answer without hedging.

  1. Get the full list of named sub-processors, the actual company names, not a category description like 'cloud infrastructure partners.'
  2. For each one, confirm exactly what data it receives: the whole document, extracted text, or only retrieved passages.
  3. Confirm whether any sub-processor uses your data to train its own models, and get that in writing, not implied by the primary vendor's own policy.
  4. Ask how a sub-processor is added in the future. You want advance notice, not a silent update to a terms page.
  5. Cross-check the list against the vendor's own description of its data flow. A mismatch between the two is a bigger red flag than either document alone.

This is the same ground covered from the buyer's side in questions to ask an AI vendor before uploading client documents, a sub-processor list is where several of those questions get answered on paper instead of in a sales call.

Why does a no-training clause matter if sub-processors still see your text?

Because a no-training promise is about use, not about access. A sub-processor can receive your document text for the narrow purpose of generating an embedding or drafting an answer, and still never use that text to improve its own models. Those are two separate commitments, and a vendor's sub-processor list should make both explicit rather than letting one imply the other. When you evaluate the clause, look for whether it names the sub-processor specifically, stating that the embeddings provider itself does not train on submitted text, rather than only describing the primary vendor's own policy.

The cleanest way to get an unambiguous answer is to ask the vendor for the sentence in writing: name the sub-processor and state, without qualification, that it does not train on data submitted through the product. A vendor that can produce that exact sentence has clearly asked its own sub-processors the same question you are asking it.

If your firm handles work under an ethical duty of confidentiality, treat every product data flow as if it will eventually be described in an engagement letter or a client audit. The tools that hold up under that description are the ones whose sub-processor list already matches what actually happens in production, not just what the marketing page says. Vorticel's own documentation describes this same data flow at the field level, for teams that want to check it against their own client-confidentiality obligations before approving the tool.

The most common objection at this stage is reasonable: your firm already vets vendors itself, so why does this vendor's list matter more than your own review? It does not replace your review, it is the input your review needs. A firm's own vendor-diligence process is only as good as the disclosures it has to work with, and a vendor that names its sub-processors and states the minimum data each one receives has given your reviewer something concrete to sign off on, instead of a policy paragraph to take on faith.

How do you start checking a vendor's sub-processor list today?

Ask for the list in writing before you upload a single document, not after. If you are evaluating Vorticel specifically, its full data flow and both named sub-processors are on the security page today, in the same detail shown above. Start a free trial and read that page side by side with your firm's own document-handling policy. It takes a few minutes and tells you, before any client file is involved, exactly who would see what.

Frequently asked questions

What is a sub-processor in plain English?

A sub-processor is any outside company your primary vendor sends your data to as part of delivering its service. If a legal AI tool uses a separate company to generate embeddings and a separate company to draft answers, both of those companies are sub-processors, even though you never signed a contract with them directly.

Is a sub-processor list the same as a security page?

No. A security page describes encryption, access control, and tenant isolation. A sub-processor list is narrower: it names each outside company that receives your data, what they receive, and why. Some vendors bury this in a security page's fine print; a vendor confident in its answer usually puts it in a visible table.

Why does it matter if a vendor uses two sub-processors instead of one?

Every sub-processor is a place your data can be handled with different terms than your contract with the primary vendor. Two sub-processors mean two sets of data-handling terms to check, not automatically more risk — what matters is whether each one receives only the minimum data its step requires, not your entire document.

Does 'no training on your data' cover sub-processors too?

It should, and you should ask directly rather than assume. A vendor can honestly say it does not train on your data while a sub-processor's default terms allow exactly that. The clause you want is the sub-processor's own commitment, not just the primary vendor's marketing language.

What should I ask if a vendor's sub-processor list is missing or vague?

Ask for the list directly and in writing: which companies, what data each receives, and their data-handling terms. A vendor that cannot produce this on request, or that describes it only as industry-standard AI providers without naming them, has not actually done the diligence it is implying.

Try grounded, cited answers on your own documents.

Create your firm's private workspace, upload a few files, and ask the first question in minutes. 14-day free trial, no card required.