Agentic legal work when you are the regulated one
Decide which legal and compliance work an agent may draft, which a licensed human must certify, and what record you keep so a regulator can follow it.
on this page · 0 / 0 checked
The memo comes back before you have finished reading the email that asked for it. It has headings, it cites the rule by number, it uses the phrase your regulator uses, and it reads better than the one you would have written on a Thursday afternoon. That is the problem. A draft that is fluent everywhere is harder to check than a draft that is obviously thin, because the part that is wrong is a rule number, a filing deadline, or a case that does not exist, sitting inside a paragraph that sounds correct.
This guide is for a small regulated operation: an advice firm, a broker, a clinic, a bookkeeping practice, a solo lawyer, anyone whose work product is checked by somebody with the power to fine them. It is about where the line sits between work an agent may produce and work a licensed human has to certify, and about the record that makes the difference visible afterwards. It is not for a firm that already has a compliance department and outside counsel on retainer, who have their own process for this. It is also not an answer to whether AI can replace your lawyer. It cannot, and the more interesting question is what it can hand your lawyer.
The vendor terms already made part of this decision for you
Before any regulator reaches you, your own supplier has a view. Anthropic’s Usage Policy, effective 15 September 2025, names a set of high-risk domains that includes “legal interpretation, legal guidance, or decisions with legal implications”, “healthcare decisions, medical diagnosis, patient care, therapy, mental health, or other medical guidance”, insurance “underwriting, claims processing, or coverage decisions”, and “financial decisions, including investment advice, loan approvals, and determining financial eligibility or creditworthiness” [1]. For those uses it sets two conditions. The first is human review: when you use the products to provide advice, recommendations, or in subjective decision-making directly affecting individuals or consumers, “a qualified professional in that field must review the content or decision prior to dissemination or finalization” [1]. The second is disclosure: if outputs are presented directly to individuals or consumers, “you must disclose to them that you are using AI to help produce your advice, decisions, or recommendations”, and a consumer-facing chatbot must disclose to users that they are interacting with AI rather than a human, at a minimum at the beginning of each chat session [1].
OpenAI’s usage policies, effective 29 October 2025, land in the same place by a different route. They prohibit the “provision of tailored advice that requires a license, such as legal or medical advice, without appropriate involvement by a licensed professional”, and separately prohibit the “automation of high-stakes decisions in sensitive areas without human review”, listing legal, medical, insurance, financial activities and credit, employment and housing among those areas [2].
Read those as commercial terms rather than as ethics. The word doing the work in both is the qualifier: a qualified professional, appropriate involvement by a licensed professional. Neither policy forbids the agent from drafting. Both forbid the draft from being the finished thing. If you run an agent that answers client questions about their policy coverage or their tax position with no licensed person in the chain, you have a regulatory exposure and an account-termination exposure at the same time, and only one of those comes with a warning email.
A technology-neutral rulebook already covers your agent
The common hope in a small regulated business is that the rules have not caught up yet. They have, by not moving. FINRA’s 2026 Annual Regulatory Oversight Report states that its rules, “which are intended to be technologically neutral”, and the securities laws more generally, “continue to apply when firms use GenAI or similar technologies”, and that such use “can implicate rules regarding supervision, communications, recordkeeping and fair dealing” [3]. Where a firm leans on these tools inside its own supervisory system, the report says its policies and procedures “may consider the integrity, reliability and accuracy of the AI model” [3]. On agentic systems specifically, FINRA names agents “acting autonomously without human validation and approval”, agents that “may act beyond the user’s actual or intended scope and authority”, multi-step reasoning that makes outcomes “difficult to trace or explain, complicating auditability”, and agents on sensitive data that “may unintentionally store, explore, disclose or misuse sensitive or proprietary information” [3].
The lawyer side reached the same conclusion two years earlier. ABA Formal Opinion 512, issued 29 July 2024, was the association’s first formal opinion covering the use of generative AI in the practice of law, and it did not create a new duty [4]. It restated four old ones: competence under Model Rule 1.1, which requires understanding “the benefits and risks associated” with the technology used to deliver legal services; confidentiality under Rule 1.6; communication with the client under Rule 1.4; and fees under Rule 1.5 [4].
That pattern is the durable lesson, and it survives every model release. Nobody has to write an AI rule for your AI to be regulated, because the obligation attaches to the outcome, not to the instrument that produced it. Advice is advice whether it came from a partner, a paralegal or an agent. The suitability requirement does not soften because a model drafted the recommendation. If you are looking for the line, it is not “may I use this”. It is “who is answerable for what came out, and can they show their work”.
The failure that keeps reaching courtrooms is a citation that does not exist
There is now a public count of what happens when nobody checks. The AI Hallucination Cases database, which tracks decisions where “the use of AI, whether established or merely alleged, is addressed in more than a passing reference by the court or tribunal”, listed 2,009 cases as of 3 September 2026, with 1,378 of them in the United States [5]. Lawyers accounted for 799 of them and self-represented litigants for 1,156 [5]. Recorded monetary penalties in the set run from 1 US dollar to 14,500 [5]. The maintainer is explicit about the scope: it “does not cover mere allegations of hallucinations, but only cases where the court or tribunal has explicitly found (or implied) that a party relied on hallucinated content or material” [5]. It is the decisions where a judge said so on the record, not the wider universe of every fabricated citation.
Take the shape of the failure rather than the total. What a court found in each of those decisions is reliance on material that was not real [5], and that defect sits at the level of a single reference inside a filing that otherwise reads as competent work. Which means verification cannot be done by reading. It has to be done per reference, against the source, by someone who knows what the source should say. A person who skims an agent’s output and thinks that it reads right has performed no check at all.
Buy or build the check, not the draft
The interesting movement in this market is not better drafting. It is who signs. Norm Ai raised 120 million dollars at a 1.2 billion dollar valuation on 7 July 2026, led by Khosla Ventures, on a structure it calls agentic law: an affiliated firm, “Norm Law, LLP, an affiliated AI-native law firm running on the Norm Ai platform”, which “uses those AI agents to serve clients as outside counsel, with senior attorneys supervising, calibrating, and improving the agents”, and which “prices based on outcomes rather than hours” [6]. Whatever happens to that company, the shape is worth copying, because the product being sold is not the draft. It is the supervision around the draft, and a named professional at the end of it.
You can build a small version of this yourself, and it has three parts. The agent produces the work. The work is checked against a stated standard, meaning an actual document you can point at: the rule text, the policy wording, the client’s file, your own house checklist, not the phrase “for accuracy”. A licensed or otherwise accountable human then certifies the check, by name, before anything leaves the building.
Most small operators skip the middle part, and the middle part is the whole thing. “I reviewed it” is not a control. “I compared each rule citation against the current text on the regulator’s site, and the two figures against the client’s statement” is a control, because it can be repeated by someone else and can fail visibly. Write the standard down once per work type. Contract review gets one, client suitability letters get another, complaint responses get a third.
Confidentiality is a settings problem before it is a judgment problem
Formal Opinion 512 puts the confidentiality duty plainly: a lawyer using these tools “must be cognizant of the duty to keep confidential all information relating to the representation of a client, regardless of its source, unless the client gives informed consent” [4]. The same instinct applies to a client’s financial statements, an employee’s medical note, or a customer’s identity documents, whether or not a bar association is watching. The practical consequence is that the account you use matters as much as the prompt.
The vendors publish the settings. OpenAI states that it does not train its models on your data by default across ChatGPT Business, ChatGPT Enterprise, ChatGPT Edu and the API platform; that it “may securely retain API inputs and outputs for up to 30 days”, with zero data retention available for eligible endpoints and a qualifying use case; that it can execute a Data Processing Addendum for ChatGPT Business, ChatGPT Enterprise and the API; and that it is “able to sign Business Associate Agreements (BAA)” for HIPAA compliance on the API platform, on request [7]. Anthropic says that some Claude API and Claude Code for Enterprise customers, “subject to Anthropic’s approval”, may have arrangements under which it does not store their inputs or outputs, and that its Business Associate Agreement is for customers of its HIPAA-eligible services, who “are subject to certain configuration requirements and limitations on what features/integrations are available” [8].
Two things follow. First, the personal plan you use for everything else is not the account regulated material goes into, and the fix is administrative rather than technical. Second, the agreement you need has a name, so ask for it by that name and keep the signed copy. “They said it was fine” is not a record. A countersigned BAA is.
The record is the deliverable
FINRA’s report is unusually concrete about what to keep: “storing prompt and output logs for accountability and troubleshooting; tracking which model version was used and when; and validation and human-in-the-loop review of model outputs” [3]. That is a small firm’s whole compliance artefact for AI, written by the regulator.
At your scale it is one table, one row per piece of regulated work: the date, the tool and model version, the task or prompt, the standard it was checked against, who reviewed it, what they changed, and when it went to the client. Fill the row while the work is still open, because nobody reconstructs it later. It costs almost nothing until the day someone asks how a document was produced, at which point it is either the cheapest thing you ever did or the reason you are rebuilding six months of work from memory.
Billing deserves the same discipline. Opinion 512 permits a lawyer to bill for the time spent inputting information into the tool and reviewing the resulting output, but says that “in most circumstances, the lawyer cannot charge a client for learning how to work a GAI tool” [4]. Read that as the general rule for anyone who bills time: the review is real work and is billable, the speed gain belongs to the client, and the learning curve is yours.
documents × review minutes ÷ 60 × hourly cost. Computed in the page; nothing is sent anywhere.
That number is the real price of agentic legal work, and it is the one that never appears in a demo. If it is larger than what the work costs you today, the agent is not saving you anything, and the honest move is to narrow the work type until the check is short enough to be worth doing properly.
What still goes wrong
Review theatre is the main failure, and it gets worse as the drafts get better. A reviewer who has approved a long run of clean outputs stops opening the source document, because the base rate has taught them the check is a formality. Lawyers, whose training is built around checking citations, still account for 799 of the recorded decisions [5]. Build the check so that it produces evidence, a marked-up copy or a filled checklist, rather than a feeling, and rotate the reviewer if you can.
The ground also moves under a control that was correct when you wrote it. The model behind your approved prompt is replaced, the usage policy is updated on a date you did not notice, the retention setting you chose applies to one product and not the one you migrated to. Both vendor policies cited here carry 2025 effective dates and will carry later ones [1][2]. A control with no review date is a control that is quietly expiring, so put the vendor policy pages and your own standards on the same quarterly pass.
Finally, nobody can tell you exactly what adequate supervision of an agent looks like in your sector yet. FINRA names auditability as a complication rather than a solved problem [3], and the bar guidance restates old duties rather than mapping them onto autonomous systems [4]. You are making a judgement call in a gap. The record is what turns that judgement from a defence you improvise later into one you can show, which is the only reason to keep it.
- 01Anthropic — Usage Policyanthropic.com
- 02OpenAI — Usage policiesopenai.com
- 03FINRA — 2026 Annual Regulatory Oversight Report, Gen AIfinra.org
- 04American Bar Association — ABA issues first ethics guidance on a lawyer's use of AI tools (Formal Opinion 512)americanbar.org
- 05Damien Charlotin — AI Hallucination Cases databasedamiencharlotin.com
- 06Norm Ai — Series C announcement (PR Newswire)prnewswire.com
- 07OpenAI — Enterprise privacyopenai.com
- 08Anthropic Privacy Center — Shipping a product with Claude that handles personal dataprivacy.claude.com