friday, september 18, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

How to tell whether AI is actually working for you

Three numbers you can collect yourself in four weeks, why your impression of the time saved is unreliable, and how to read the adoption figures other people publish.

Published 2026-09-04 · Updated 2026-09-04 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

You pay for two or three AI subscriptions. Someone asks whether they are worth it, and the honest answer is a shrug and a story about the afternoon the thing wrote a proposal in eight minutes. That story is real. It is also the same class of evidence that vendor case studies run on, gathered from a smaller sample, by a person with a reason to like the result. It will not tell you whether to renew.

This guide is about the small set of numbers that survive contact with a real business, and about the two reliable ways the obvious numbers mislead. It is written for a one-person business or a team under about twenty people spending its own money, where nobody is going to build a benefit-realisation model and the decision is simply whether to keep paying. If you have a procurement function and a finance partner, you need a heavier document than this one. If you have four seats and a nagging feeling, you are in the right place.

The headline number is the one that was easiest to produce

Every published AI statistic was chosen by someone, and the selection is not random. Reach is the cheapest thing to measure, so reach is what gets published. OpenAI’s own study of consumer usage, based on about 1.5 million conversations, reported 700 million weekly active users of ChatGPT, and also found that roughly 70% of consumer messages were not work related [5]. Both numbers are from the same organisation, in the same document, and only one of them tends to travel.

Total users and total interactions measure how many people came into contact with a tool. They say nothing about whether anyone got value from it, because a product that 10,000 people tried once and abandoned produces a larger number than one that 500 people depend on daily, and the two look identical in a press release that stops at the top line. The same trap operates at your scale. “Everyone on the team has an account” is a reach number. It is the small-business version of a weekly active user count, and it is worth about as much.

The base rate is also lower than the conversation implies. When the Census Bureau measured AI use directly through its Business Trends and Outlook Survey, the bi-weekly estimate of firms reporting AI use for business purposes rose from 3.7% to 5.4% between September 2023 and February 2024, with an expected rate of about 6.6% by early autumn 2024, and AI use turned out to be higher in large firms while the relationship with firm size was not monotonic [3]. Those readings stop in February 2024, so read the level as history rather than as a current count. The shape is what carries. Reported use ran far below the volume of noise about it, and it was not a simple function of how big you were. On those numbers nobody was falling behind at the rate the marketing suggested, and you can afford to measure before you commit.

Activation and weekly use are the adoption numbers worth keeping

Two numbers do carry information, and the better vendor case studies publish them. OpenAI’s write-up of Univé, one of the Netherlands’ largest cooperative insurers, published on 31 July 2026, reports that 97% of its ChatGPT Enterprise licences were activated, that 85% of licensed users were active every week, and that employees averaged 40 prompts per active user each week [6].

Read what each one actually measures. Activation is close to an IT event. It says the licence was provisioned and the person logged in at least once, which any competent rollout can achieve by sending an email with a deadline in it. The weekly figure is the harder one, because it has to survive the novelty wearing off, and a rate that high months into a rollout usually means the tool sits inside work people were doing anyway rather than beside it. Prompts per active user is the sanity check on the other two: it separates people who open the tab from people who use the thing.

Your version of this takes ten minutes a month and no dashboard. Count the seats you pay for. Count how many of them were used in the last seven days, not at any point since signup. Write both numbers in the same file every month, on the same day. That is your activation and weekly-active rate, and it is the only adoption data about your business that exists.

The gap between those two counts has a price attached. A Claude Team standard seat is $25 per seat per month billed monthly and $20 per seat per month billed annually [7]. ChatGPT Business is priced on the same shape, $25 per user per month monthly and $20 per user per month annually, with a workspace minimum of two seats [8]. Three unused seats out of eight is $60 to $75 a month, which is $720 to $900 a year, spent on people who tried the tool in week one and went back to what they were doing. Both plans bill by the seat rather than by use [7][8], so the invoice looks identical whether the seat was opened or not, and the cost stays invisible until someone counts.

Your impression of the time saved is not evidence

The uncomfortable finding is that people who use these tools heavily are bad at estimating what the tools did for them. METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real issues from their own repositories, averaging about two hours each, randomly assigning each issue to allow or disallow AI [1]. The developers were not novices, with “dozens to hundreds of hours” of prior experience prompting language models [1].

When they were allowed to use AI tools, they took 19% longer to complete issues [1]. Before starting, they had expected AI to speed them up by 24%. After finishing, having actually been slower, they still believed AI had sped them up by 20% [1]. The direction of the error is the point. The estimate was not noisy, it was confidently wrong in the flattering direction, and it stayed wrong after the experience that should have corrected it.

METR is careful about the limits of this, and you should be too. The write-up lists what the result does not show: that AI fails to speed up many or most software developers, that it fails in domains other than software development, or that there is no more effective way of using the same systems in the same setting [1]. The authors also say they cannot rule out learning effects beyond 50 hours of Cursor use [1]. The finding does not mean AI made you slower. It means that the feeling of having been made faster is produced whether or not you were, so the feeling cannot be the evidence. If you are going to keep paying, keep paying on the strength of a number you wrote down.

A four-week measurement that costs about an hour

Pick one task. It has to be recurring, it has to have an end you can recognise, and it has to be work you were going to do regardless. Weekly client update. Monthly invoice reconciliation. First draft of a proposal. Vague candidates like “research” do not work, because you cannot tell when they are finished.

Write down how many times that task happens in a month. Then time it five times the way you do it now, with a phone timer, to the nearest minute, and stop the clock when the work is genuinely done rather than when the first draft appears. Five is a small sample and you should treat the result as an estimate, but five recorded times beat one remembered impression by a wide margin.

Then do the same task five times with the tool, and time it the same way. Count the whole loop, including the prompt you rewrote twice and the two paragraphs you fixed afterwards. The most common way a small measurement inflates itself is by starting the clock when the output appears and ignoring the review. If the tool produces something that needs checking, checking is part of the task.

Subtract, multiply by frequency, and you have monthly hours. Then decide what an hour is worth to you, which is either your billing rate if the hour is billable or a number you can defend if it is not, and compare it to the subscription.

calculator
What one tool is worth per month
$ / month

minutes saved per week × 4.33 weeks ÷ 60, times your hourly rate, minus the monthly cost. A negative result means the tool is not paying for itself on this task. Computed in the page; nothing is sent anywhere.

Two rules keep this honest. Record at least one case where the tool made the work worse or longer, because a run of five with no failures usually means you quietly excluded one. And do not run the measurement on the demo task the vendor suggests, which is selected to work.

What a case study proves, and what the 95% figure proves

Set the two most-quoted kinds of AI statistic next to each other. Univé’s rollout shows 97% activation and 85% weekly use [6]. The GenAI Divide report from MIT’s NANDA initiative, surveying enterprise deployments, reported that 95% of organisations are getting zero return despite $30 billion to $40 billion of enterprise investment in generative AI, and that just 5% of integrated pilots are extracting millions in value while the vast majority sit with no measurable profit-and-loss impact [2].

Both can be true, because they measure different things. Activation and weekly use are adoption metrics. Return is a profit-and-loss metric. A company can genuinely have most of its staff using a tool every week and still not be able to point at a line in its accounts that moved, and that is the normal outcome rather than a scandal. For you the practical consequence is that the two questions have to be asked separately and in order. First, whether anyone is using it. Second, what stopped being done by hand.

The second question is where the value actually lives. Anthropic’s Economic Index found that 77% of its first-party API transcripts showed automation patterns, especially full task delegation, against 12% for augmentation, while in a sample of Claude.ai conversations the split between automation and augmentation was nearly even, with directive conversations rising from 27% in late 2024 to 39% in the August 2025 sample [4]. Programmatic use lends itself to handing over a whole task, and chat lends itself to helping with one. Both are legitimate, but only the first tends to remove a recurring cost, because a task that leaves your desk entirely stops consuming your attention while a task you did with help still occupies the same slot in your week.

So when you read any AI announcement, including this one, do three things. Find out which number is being reported and what it would look like if the deployment had failed. Check whether anyone measured after the novelty period. And separate the usage claim from the money claim, because those are almost never the same claim.

checklist
Before you decide a tool is worth renewing
0 of 8 · saved in this browser only

What still goes wrong

The measurement is small and you should hold it loosely. Five timings on a task that varies in size will produce a number with real uncertainty around it, and you cannot fix that without spending more time measuring than the tool is likely to save. The correct response is not more rigour, it is a wider margin. If the tool has to save 40 minutes a week to break even and it saved 45, that is a coin flip, not a result. If it saved four hours, you have your answer and further precision is a waste of your afternoon.

The deeper problem is that the things easiest to time are rarely the things that matter most. Minutes on a repeated task are countable. Whether a proposal was more persuasive, whether the code you shipped will hold up in eight months, whether the summary quietly dropped the one detail the client cared about, none of these appear in a stopwatch, and some of them only surface as costs much later. A tool can pass this test convincingly and still be a bad idea, which is why the checklist asks you to log the failures too. Quality damage tends to arrive as an anecdote long before it arrives as a metric, and an anecdote about harm deserves more weight than an anecdote about speed, because nobody is motivated to invent it.

Finally, a positive number is not the end of the argument. A tool that saves three hours a month on a task you do not care about has still cost you the hours you spent learning it and the attention you spend maintaining it. And the numbers you collect describe the tool as it is configured today, on models that will be replaced. Re-run the count when something significant changes, treat the result as a snapshot with a date on it, and be as willing to act on a negative number as you were hoping to act on a positive one. Cancelling is a valid outcome of a measurement, and it is the outcome nobody plans for.

sources
  1. 01METR — Measuring the impact of early-2025 AI on experienced open-source developer productivitymetr.org
  2. 02MIT NANDA — The GenAI Divide: State of AI in Business 2025 (report PDF)mlq.ai
  3. 03US Census Bureau — Tracking Firm Use of AI in Real Time (CES-WP-24-16)census.gov
  4. 04Anthropic — Anthropic Economic Index report, September 2025anthropic.com
  5. 05OpenAI — How people are using ChatGPTopenai.com
  6. 06OpenAI — Univé builds an AI-ready workforceopenai.com
  7. 07Anthropic — Claude pricingclaude.com
  8. 08OpenAI Help Center — What is ChatGPT Businesshelp.openai.com
next guide
The AI data shortage, and who gets to train on your text
9 min · verified 2026-09-04
related guides