legal AI · practice management · verification

Where AI Actually Helps in a Law Practice, and Where It Doesn't

An honest division of legal work into the parts machines handle well, the parts they handle badly, and the discipline that separates the two.

5 min read

Most writing about AI in law is either a sales pitch or a warning. Neither is much use when you are deciding whether to put a tool into your actual practice this month. What follows is an attempt at the more boring and more useful thing: a division of legal work into the parts machines currently handle well, the parts they handle badly, and how to tell which is which.

The shape of a task machines are good at

The tasks where AI reliably helps share a shape. They involve a large volume of material, a well-defined thing to find in it, and an answer that can be checked against a source.

Searching a body of documents for every mention of a term. Segmenting a long transcript by topic. Finding the moment in hours of recorded footage where something specific was said. Producing a first-pass chronology from material that is not in chronological order. These are all retrieval and structuring problems dressed up as reading problems, and the volume that makes them punishing for a person is nearly free for a machine.

Notice what these have in common: for each of them, you can verify the answer. The tool says the witness discussed the contract at a particular page and line; you look, and either they did or they did not. That property is what makes the task safe to delegate.

The shape of a task machines are bad at

The mirror image is equally consistent. Machines do badly where the work requires judgment, where the answer cannot be checked against a source, and where somebody has to be accountable for being wrong.

Deciding which of two arguments will land with a particular judge. Assessing whether a witness is credible. Determining whether a novel set of facts fits a doctrine. Advising a client on risk. These are not harder versions of retrieval problems; they are a different kind of problem, and the fact that a language model will produce fluent output about them is not evidence it is doing them.

The failure mode to watch is when a judgment task is disguised as a retrieval task. "Find the cases supporting this position" looks like search. It is not — it embeds a judgment about what supports the position, and that is where things go wrong.

The case everyone cites, and what it actually teaches

In Mata v. Avianca, decided in the Southern District of New York in June 2023, two attorneys and their firm were sanctioned after filing a brief containing six fabricated case citations produced by ChatGPT. The sanction was five thousand dollars, but the professional cost was considerably higher, and the case is now the standard reference point in every conversation about AI in litigation.

The lesson usually drawn from it is "do not use AI." That is the wrong lesson, and it is worth being precise about why.

The tool generated citations that did not exist. That is a known and well-documented behaviour of language models asked to produce references from memory. What turned a known limitation into sanctions was that the output was filed without anyone checking whether the cases were real — and then, when opposing counsel could not locate them, the response was to supply further unverified material rather than to check.

The failure was not the use of a tool. It was the absence of verification at the point where verification was the entire job.

The discipline that separates them

Which suggests a single practical test for any AI tool entering a practice: can you check its output against a source, in one step, without redoing the work?

If a tool tells you a witness said something and gives you the page and line, you can check it in seconds. If it tells you footage shows something and gives you the timestamp, likewise. If it tells you a case supports your position and gives you a citation, you can pull the case — and you must, every time.

If a tool produces an assertion with nothing to check it against, you have not saved any work. You have moved the work to a place where you are less likely to do it.

This is why citation is not a nice feature in legal AI. It is the mechanism that makes delegation safe, and a tool without it is asking for a trust it has not earned.

What this looks like in practice

A few things that follow, if you accept the framing:

Put AI on volume problems first. Document review, transcript search, footage review, first-pass chronologies. High tedium, low judgment, verifiable output. The return is immediate and the downside is bounded.

Be suspicious of anything that writes for you without sources. Drafting assistance is genuinely useful, but a draft containing assertions of law that you did not verify is a liability with your name on the signature block.

Test on material you already know. Run a tool over a matter you have worked. You know what is in it. Whether the tool finds what you know is there — and whether its references resolve correctly — tells you more than any vendor demonstration.

Keep the accountability where it belongs. No court will accept that a tool made the error. That is not a limitation of current technology; it is a feature of professional responsibility, and it will not change when the models improve.

The honest summary

AI is very good at the part of legal work that consists of finding things in large amounts of material, and not good at the part that consists of deciding what to do about them. Most of the value available today is in the first category, most of the risk is in confusing the second for the first, and the practice that keeps them separate is insisting that every output can be checked against something real.


This post is general commentary and is not legal advice. Professional responsibility obligations regarding the use of technology vary by jurisdiction.