How to run a two-week AI pilot at a small law firm

In shortA useful AI pilot at a small firm fits in two weeks: choose one closed matter with documents you know well, define three test questions including one the files cannot answer, have your compliance reviewer read the vendor's security page in week one, and measure minutes saved per question in week two. Decide on evidence, not on a demo.

A small firm can run a decisive AI pilot in two weeks by choosing one closed matter, writing three test questions in advance, having the firm's vendor reviewer read the security documentation in the first week, and measuring minutes saved on real questions in the second. The result is a yes or no backed by your own documents, not by a vendor's demo.

An AI pilot is a short, bounded trial of a tool on your own material, with pass criteria written down before you start. The bounded part matters. Pilots that begin as a general exploration tend to end as a shrug: nobody remembers what would have counted as success, so nobody can say whether it happened.

The pain that prompts a pilot is usually specific. An associate is spending hours re-reading a discovery production to answer the same handful of questions. A partner wants to know, across a shelf of leases, which ones carry a particular indemnity, and the honest answer is that finding out takes a day. Meanwhile someone on staff is already pasting clauses into a consumer chatbot, and nobody signed off on that.

What should you decide before day one?

Three things, all written down. First, the matter: a closed one whose documents you know well, so you can grade answers. Second, the participants: two or three people, including whoever reviews vendors. Third, the pass criteria. A workable set is that every answer cites a document and page you can open, the tool refuses at least once when it should, the reviewer's confidentiality questions are answered in writing, and week-two questions save a meaningful amount of time.

Write your three test questions now, before you see the tool:

  • One question with a clear answer in a specific document, so you can check the citation.
  • One question whose answer is spread across several documents, to test retrieval breadth.
  • One question you know the documents do not answer, to test whether the tool refuses or guesses.

How does week one go?

Week one is about correctness and confidentiality, not speed. The day-by-day plan below assumes a tool with a free trial and no card, which removes the procurement step entirely.

  1. Day 1: create the firm's account, invite the two other participants, and upload the closed matter's documents. Confirm each one reaches a ready state.
  2. Day 2: ask the three prepared questions. For each, open the cited passage and mark the answer as correct, partly correct, or wrong. Record whether the unanswerable question drew a refusal.
  3. Day 3: the reviewer reads the vendor's security page and sub-processor list and sends the vendor any question that is not answered there. Written questions, written answers.
  4. Day 4: each participant asks five questions of their own choosing about the matter and grades them the same way.
  5. Day 5: a ten-minute meeting. Tally citations opened, answers graded, refusals observed, and the reviewer's open items. If a pass criterion is already failing, stop here and save a week.

The step that firms most often skip is day three. Do not skip it. A tool that answers well but cannot document where client text goes is not a tool you can adopt, and finding that out in week one is the whole point of the schedule.

How does week two change?

Week two is about value on live work, and it only starts if the reviewer signed off. Add one active matter, with the responsible attorney's agreement, and switch from grading to measuring.

Each participant answers five real questions during the week and notes two numbers per question: minutes the answer would have taken by hand, and minutes it took with the tool including the time to open and verify the citation. Round to five-minute increments. This is where the manual method shows its cost most clearly, because the by-hand estimate for a question like which of these agreements contain an assignment restriction is measured in hours, while a grounded tool built for this, such as Vorticel, returns the list with page citations in the time it takes to read them. The pricing page shows which plan fits the number of seats and documents you used during the pilot.

Keep asking the occasional question the documents cannot answer. Consistent refusal on live work is as important as it was on the closed matter, because live work is where a confident guess would do harm.

What does the decision meeting look like?

Thirty minutes at the end of week two, with the four pass criteria on one page. For each, the evidence is already collected: the grades from week one, the refusal count, the reviewer's written answers, and the minutes saved per question from week two.

CriterionEvidence from the pilotPass looks like
Every answer citedWeek-one gradesCitations opened and matched for all correct answers
Refuses when it shouldUnanswerable questions askedRefusal every time, no invented answers
Confidentiality documentedReviewer's written Q and ANo open items on isolation, encryption, training, sub-processors
Time savedWeek-two minute notesA clear majority of questions faster including verification

If all four pass, choose the plan and move the closed-matter documents out if you prefer a clean start. If one fails, you have a specific, documented reason, which is a better outcome than most pilots produce.

What are the common ways a pilot goes wrong?

The first is uploading too much. Forty documents from a matter you know is a better pilot than four hundred from one you do not, because you cannot grade what you cannot check. The second is letting the demo stand in for the test: a vendor's sample documents prove nothing about your files. The third is skipping the reviewer until the end, which turns a one-day check into a blocker after everyone else has already decided.

The last is measuring the wrong thing. Raw answer speed is not the metric; time to a verified answer is. A tool that answers in two seconds but sends you re-reading the document to confirm has saved nothing. A citation you can open is what makes the second number small.

The objection worth taking seriously

The honest reason a small firm hesitates is not cost, it is attention. Two weeks of even light-touch testing competes with billable work. The plan above is built around that: roughly an hour per participant in week one, real work in week two that would have been done anyway, and two short meetings. If the firm cannot find that, the answer is to wait, not to adopt a tool untested.

Start the clock

Create your firm's free trial, invite the two other participants, and upload the closed matter today. The trial lasts 14 days, which is exactly the length of this plan, and it needs no card, so day one of the pilot is also day one of the trial.

Frequently asked questions

How many people should take part in a pilot?

Two or three is enough: one partner or senior associate who knows the matter, one paralegal or associate who will do most of the day-to-day asking, and whoever reviews vendors for the firm. More participants add coordination without adding evidence, and a small group can meet for ten minutes at the end of each week.

Should we use a live matter or a closed one?

Use a closed matter for the first week. You already know the answers, which is what lets you grade the tool's citations and refusals honestly. In the second week, if the reviewer has signed off, add a live matter so you can measure time saved on questions you genuinely need answered.

What if the tool refuses to answer a question we think it should?

Check whether the document that holds the answer was actually uploaded and finished processing. If it was, note the question and the passage you expected, and count it as a retrieval miss. A few misses on hard questions are normal; a refusal on a question plainly answered by an uploaded document is a defect worth reporting to the vendor.

How do we measure time saved without a stopwatch culture?

Ask each participant to note, for five questions during week two, how long the answer would have taken by hand and how long it took with the tool including verification. Round to the nearest five minutes. The goal is an honest order of magnitude, not a time study, and five data points per person is enough to see it.

What should the decision meeting cover?

Four things: whether every answer carried a citation you could open, whether the tool refused when it should have, whether the reviewer's confidentiality questions were answered in writing, and the rough minutes saved per question. If all four are satisfactory, pick a plan; if one is not, you have a specific reason to pass.

Try grounded, cited answers on your own documents.

Create your firm's private workspace, upload a few files, and ask the first question in minutes. 14-day free trial, no card required.