Questions to ask an AI vendor before uploading client documents

In shortBefore a firm uploads privileged material to any AI tool, it should get written answers to seven questions: who else can see the data, how it is encrypted at rest, whether it is used to train models, which sub-processors receive it, how deletion and export work, what is logged, and how the vendor proves its claims rather than asserting them.

Before any privileged document leaves your firm for an AI tool, you should have written answers to seven questions: who else can see the data, how it is encrypted at rest, whether it trains a model, which outside companies receive it, how deletion and export work, what gets logged, and how the vendor proves any of that. A vendor that answers all seven specifically is one you can evaluate; a vendor that answers in generalities is one you should walk away from.

Vendor diligence for AI is the process of confirming, before uploading anything, exactly what happens to a document from the moment it is uploaded to the moment it is deleted. Most firms already do this for cloud storage and practice-management software. AI tools deserve the same review, with two additions: the data is often sent to outside model providers, and the product may be tempted to answer from outside knowledge rather than from your files.

The pain is real. A partner wants faster answers from a discovery production; an associate has already pasted a paragraph of a retainer into a general chatbot; nobody is sure where that text went. Leaving the question unsettled means either banning useful tools outright or accepting an exposure nobody has actually measured.

What should you ask about isolation between firms?

Ask how the vendor guarantees that your firm's documents can never appear in another customer's results. The right answer describes a tenant boundary enforced at the data layer: every document, extracted passage, question, and answer is tagged with your firm's account, and every lookup filters on that tag before anything is returned.

Follow up with a harder question: is there any shared search index across customers? Some products build one index for everyone and rely on filtering at the interface. That is weaker than a design where no query path exists that can return another firm's material. You want the second design, and you want the vendor to say so in those words.

How is the data encrypted, and when does it leave the vendor?

Encryption in transit is table stakes. The questions that separate vendors are about encryption at rest and about which steps send data outside the vendor's own infrastructure.

A good answer names the library and the key model, for example per-firm encryption with libsodium before a file is written to disk, and states that the original file never leaves that encrypted storage. What leaves, if anything, is extracted text sent to an embedding provider so search can work, and, at question time, only the handful of passages retrieved for that question sent to a language model to draft the answer.

Here is the distinction to listen for. There is a large difference between a product that sends your whole document library to a model provider and one that sends a few retrieved passages per question. The Vorticel security page walks through this in five steps precisely because reviewers asked for it that way.

Does the vendor train on your documents?

Ask for a plain statement: are your documents, questions, or answers ever used to train any model, the vendor's own or a third party's? The answer should be an unqualified no, followed by an explanation of how the vendor chooses model providers whose data-handling terms match that commitment.

Be wary of answers that only cover the vendor's own models. If document text is sent to an embedding or language-model provider, that provider's terms matter too. The vendor should be able to tell you which terms apply and what the minimum data sent to each provider is.

Who are the sub-processors, and what does each one receive?

Every outside company that touches your data is a sub-processor, and you are entitled to a complete list. Ask for a table with four columns: the company, what it receives, when, and why the product cannot work without it.

Sub-processor typeTypical data receivedWhat to confirm
Embedding providerExtracted document text, chunk by chunk, plus question textNo retention beyond the request; no training
Language-model providerOnly the passages retrieved for one question, plus the questionNo bulk file upload; no whole-library transfer
Payment processorBilling identifiers onlyNever receives document content
Identity providerEmail identity at sign-in onlyNever receives document content
Email providerTransactional messages such as invites and resetsNo document content is ever attached

If a vendor cannot produce this table, it either does not know its own data flow or does not want you to. Neither is acceptable for privileged material.

How do deletion, export, and logging work?

Three practical questions round out the review. First, when you delete a document, does deletion remove both the encrypted file and the extracted, searchable text? Second, can you download your original files back out at any time, so your firm is never locked in? Third, what is logged: sign-ins, uploads, deletions, team changes, and, importantly, any access by the vendor's own staff?

On that last point, ask whether staff access to a customer account requires a fresh, individually verified elevation that is itself logged, and whether an impersonation feature exists anywhere in the product. The best answer is that no standing view-any-customer capability exists at all.

How does the vendor prove it rather than assert it?

Every claim above can be written into a brochure. What you want is evidence that the claims are tested. Ask what automated checks run before each release. Two are worth insisting on:

  1. An isolation check that attempts to make one firm's account retrieve another firm's documents and confirms it returns zero rows, every time.
  2. A grounding check that asks the tool a question its documents cannot answer and confirms it replies that it could not find the answer, without calling the model to guess.

A vendor that runs these checks can describe them in a sentence. A vendor that does not will change the subject.

Putting the seven questions to work

Send the questions in writing before a demo, not after. Written answers are harder to soften, and they give your compliance reviewer something to file. Score each answer as specific, vague, or missing, and treat any missing answer on isolation, encryption, or training as disqualifying.

The honest objection is cost and time: a small firm may feel it cannot afford a vendor review for every tool. The counterpoint is that the review takes an hour with a vendor that has done the work, because the answers already exist on a public page. It only becomes expensive with vendors who have not done the work, which is exactly the information you needed.

Next step

If you would rather test the answers than read them, start a free Vorticel trial, upload a few non-sensitive sample documents, and ask a question you know the files cannot answer. You will see whether the tool refuses or guesses within the first minute, and the 14-day trial needs no card.

Frequently asked questions

Is it safe to upload client documents to an AI tool at all?

It can be, but only when the vendor can show, in writing, how documents are isolated per firm, encrypted at rest, kept out of model training, and limited to named sub-processors. The risk is not the technology itself but a vendor that cannot answer these questions specifically. Treat vague answers as a no.

What does tenant isolation mean in practice?

Tenant isolation means every document, passage, question, and answer is tagged with your firm's account and every lookup is filtered to that tag before anything is returned. A properly isolated system has no shared search index across firms, so one firm's material cannot appear in another firm's results even by accident.

Why does it matter which sub-processors a vendor uses?

A sub-processor is any outside company that receives some of your data so the product can work, such as an embedding provider or a language-model provider. You need to know who they are, exactly what each one receives, and when, because your confidentiality obligations follow the data wherever it goes.

Should a firm insist on a no-training clause?

Yes. A vendor should state plainly that your documents, questions, and answers are never used to train any model, its own or a third party's, and it should be able to explain how it selects vendors whose data-handling terms match that commitment. A no-training statement that only covers the vendor's own models is incomplete.

How can a firm verify a vendor's security claims instead of taking them on faith?

Ask what automated checks run before each release. A credible vendor can describe a test that tries to make one firm's account retrieve another firm's documents and confirms zero rows, and a test that asks a question the documents cannot answer and confirms the tool refuses rather than guesses. Claims without checks are just marketing.

Try grounded, cited answers on your own documents.

Create your firm's private workspace, upload a few files, and ask the first question in minutes. 14-day free trial, no card required.