saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

Using AI without losing your edge

How to keep the judgment that lets you catch a wrong answer, while still handing the model everything that does not need you.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-05 · Rami
on this page · 0 / 0 checked

You handed over the first draft because it was faster. Then the client email, then the pricing sheet, then the scope note, then the reply to the awkward supplier. Every one of those handovers was obviously right on the day you made it. None of them came with a receipt. The cost, if there is one, shows up much later, on the day a job arrives that the tool gets subtly wrong, and you read the output twice, and it looks fine to you.

This guide is about the skill underneath the tool: which parts of your work are worth keeping in your hands, how to practise them without giving back the time you gained, and how to find out whether you have already slipped. It is written for solo operators and small teams who use AI every day and have nobody auditing their output. It is not an argument for using less AI, and it is not for someone still deciding whether to adopt it. It also will not help with a skill you never had; everything below is about experienced people losing something they used to be good at.

Skill decay from routine AI use is now a measured effect

The clearest evidence so far comes from medicine, which is inconvenient, because medicine is not your business. Take it as a direction of effect rather than a number that applies to you.

Researchers looked at four endoscopy centres in Poland between September 2021 and March 2022, covering 1,443 colonoscopies performed without AI assistance by 19 experienced endoscopists, each of whom had already done more than 2,000 procedures. Of those, 795 were performed before the centres introduced an AI detection tool and 648 after. The adenoma detection rate in the unassisted procedures fell from 28.4% before AI exposure to 22.4% after, a relative decline of about 20% [1]. One of the authors, Marcin Romańczyk, put it as plainly as a researcher can: “To our knowledge this is the first study to suggest a negative impact of regular AI use on healthcare professionals’ ability to complete a patient-relevant task in medicine” [1].

Two things about that study matter more than the headline. It is observational, and the authors say so: factors other than the introduction of AI may have influenced the result [1]. And it was restricted to experienced endoscopists, which the authors flag as a limit on how far it generalises [1]. Nobody has shown that every skill in every trade erodes at that rate. What has been shown is that a few months of handing a task to a machine was enough to move a measured outcome in people who had spent careers getting good at it.

The knowledge-work evidence is thinner but points the same way. A Microsoft Research and Carnegie Mellon survey of 319 knowledge workers, covering 936 real examples of generative AI use at work, found that the more confidence a worker had in the AI’s ability to do a task, the less critical thinking they reported applying to it, while higher confidence in their own ability went with more [2]. The paper describes the underlying shift as workers moving “from task execution to oversight” [2]. That is self-reported data, not a controlled experiment. It is still the mechanism you would expect: trust reduces scrutiny, and scrutiny is the thing you were bringing.

You will not feel it happening

The uncomfortable finding is not about skill at all. It is about your ability to judge your own performance while using these tools.

METR ran a randomised trial with 16 experienced open-source developers working on 246 real issues in repositories they knew well. With AI tools available, they took 19% longer to complete issues. Before starting, they had expected AI to speed them up by 24%. After finishing, having actually been slower, they still believed the tools had made them about 20% faster [3]. The gap between the measurement and the belief was roughly 39 percentage points, and it survived the experience.

The caveats are real and METR states them: 16 developers is small, they were experienced people in repositories they had contributed to for multiple years, the tools were early-2025, and the authors say they “can’t rule out learning effects beyond 50 hours of Cursor usage” [3]. Do not read it as “AI makes everyone slower.” Read it for the part that is hard to explain away: skilled professionals were wrong about their own throughput, in the flattering direction, while doing the work.

That is what makes deskilling hard to catch. Your sense that you have still got it is not evidence. It is generated by the same brain that is enjoying the reduced effort. The only honest test is a rep: do the task the old way, on the clock, and see.

Protect the skills where being wrong is expensive and nobody else is checking

Not every skill deserves defending. If your mental arithmetic has gone soft because a spreadsheet does it, the cost of that on a bad day is a number you check twice. If you can no longer read a paper map, the cost is a longer drive. The question is not whether a skill is fading but what happens on the day it is gone and the tool is wrong.

Three tests, applied to each task you have handed over. First, what an error costs: a typo in a newsletter is not a mispriced quote. Second, who else looks at it: in a hospital there are audits, credentialing and colleagues; in a two-person business there is you, and you are busy. Third, whether the skill is the product: if a client is paying for your judgment on their contract or your voice in their copy, that skill is the thing being sold, not an input you can outsource without telling them.

The vendors have drawn the same line for their own reasons. OpenAI’s usage policies, effective 29 October 2025, prohibit “automation of high-stakes decisions in sensitive areas without human review,” and list those areas as including legal, medical, financial, insurance, employment, housing and critical infrastructure [5]. That rule assumes something it does not say out loud. A review is only a control if the reviewer could have done the work. Otherwise the human in the loop is a signature, and the process has a person in it without having any judgment in it.

Draft the high-stakes version yourself, then bring in the model

The cheapest habit that works is an ordering rule, and it applies to a very short list of tasks: the ones that survived the three tests above.

On those, write the first pass yourself before you open the chat. Rough is fine. Then hand the draft over and use the model as a second reader: find what is missing, argue the other side, tighten the second half, check the numbers against the source. You keep the part where the thinking happens and you still get the speed on the pass where it is mostly labour. On everything else, start with the model, because most work is not the work you are being paid for.

This has a side effect worth having. When you have written your own version first, you notice disagreements. When you start from the model’s version, you get anchored to it, and the edits you make are decoration on a structure you never chose. The difference is invisible in the finished document and enormous in what you learned by producing it.

Book manual reps the way pilots do

Aviation solved a version of this problem before anyone had heard of a language model, and the regulator’s answer is instructive precisely because it is not a ban on automation.

The FAA’s Safety Alert for Operators 17007 exists to encourage “the development of training and line-operations policies which will ensure that proficiency in manual flight operations is developed and maintained for air carrier pilots” [4]. Its recommendation to operators is “Encouragement to manually fly the aircraft when conditions permit, including at least periodically, the entire departure and arrival phases, and potentially the entire flight, if/when practicable and permissible” [4]. It also warns operators away from blanket rules, telling them to avoid “overly general statements” in either direction and to leave the call to the pilot in command’s judgment about conditions on the day [4].

Translate that into your week. Periodically, not constantly. Whole phases, not fragments, because doing the easy half by hand teaches you nothing about whether you can still do the hard half. When conditions permit, meaning on the low-stakes job with slack in the deadline rather than the one that is on fire. And on your own judgment rather than a quota you will resent and abandon by March. Three or four full manual reps a quarter on each protected skill is enough to notice a decline while it is still small.

The signal you are looking for is not the finished quality. It is friction. If a thing that used to take 20 minutes now takes 45, or you find yourself reaching for the tool halfway through because you cannot remember how a section normally goes, you have your answer, and you got it on a job that did not matter.

Make the model teach rather than finish

The tools now ship modes built for exactly this, aimed at students, useful to anyone. OpenAI added study mode to ChatGPT on 29 July 2025 for Free, Plus, Pro and Team users, describing it as “a learning experience that helps you work through problems step by step instead of just getting an answer,” built on Socratic questioning, hints and self-reflection prompts “instead of providing answers outright” [6]. Google shipped Guided Learning in the Gemini app on the same principle: it “can lay out a study plan for a topic before walking you through each step,” and checks as it goes “to see if you’re picking things up along the way” [7]. Its product manager, Nupur Jain, gives the reason as “knowing how to arrive at an answer is more critical than the answer itself” [7].

You do not need the mode to get the effect, and outside of learning a subject from scratch you often do not want it. What you want is one extra request on work you were going to accept anyway. Ask for the reasoning behind a structural choice. Ask what it considered and rejected. Ask for the rule that explains an edit, in a form you could apply yourself next time. Ask it to mark the two weakest paragraphs and say why they are weak.

Each of those takes under a minute and converts a finished artefact into something you read as a practitioner rather than as an approver. It is the difference between the tool doing reps and you doing reps. It also gives you something to judge. An explanation that restates the output in different words, or cites a rule it cannot name, is a reason to read the thing again rather than approve it.

Treat full automation of a task as a decision with a date on it

The handover you should worry about is never announced. It happens because the tool was open and the deadline was close, and then it happens again, and after the fourth time it is simply how that job gets done now.

Make it explicit instead. Keep a short list of the tasks you have fully handed over, with the month you handed each one over, and a note on whether you are still able to do it. Review the list when you review anything else quarterly. Some entries will be fine and should stay automated forever. One or two will be tasks that quietly migrated from the “not my product” pile into the “actually this is my product” pile without ever being reassessed.

The pressure to keep migrating is only going up, because the tools keep getting better and stay cheap relative to an hour of your time. Anthropic currently sells Claude Pro at $20 a month billed monthly and a standard Team seat at $25, listing Opus 5 as “ideal for complex agentic coding and enterprise work” and Sonnet 5 as its “high-performance model for coding and agents” [8]. At those prices, on any given Tuesday, the case for doing a task yourself will almost never win on economics alone. It has to win on something else, which is why the decision has to be made deliberately, in advance, about a specific short list, rather than in the moment against a deadline.

checklist
Keeping your hand in
0 of 7 · saved in this browser only
calculator
Time to keep your hand in
h / quarter

skills × minutes × reps ÷ 60. Computed in the page; nothing is sent anywhere.

What still goes wrong

The evidence base is thin and none of it is about you. The colonoscopy result is observational, from four centres, in one specialty, and the authors themselves note that other factors may have driven it [1]. The METR trial is 16 developers on early-2025 tools with limited hours on them [3]. The Microsoft and Carnegie Mellon findings are self-reported [2]. Together they establish that a real effect exists in a measurable direction. They do not tell you your personal rate of decay, which skills of yours are most exposed, or how many hours of practice would hold a given skill steady. Anyone quoting you a dose is making it up.

Manual reps also cost exactly what they look like they cost. The hours in the calculator above are hours the tool was supposed to save you, and on a bad month they are the first thing to be cancelled. That is a legitimate trade rather than a failure of discipline, as long as you know you are making it. The failure is cancelling them silently for a year and then discovering the position you are in during a job that mattered.

And staying sharp fixes only the failures that were ever about your skill. It does nothing about a model that invents a citation, a price that changed last week, or an agent that follows instructions hidden in a document it was asked to read. Those need verification habits and blast-radius limits, not practice. Keeping your edge is what makes the verification worth doing, because it decides whether you notice anything when you look. It is not a substitute for looking.

sources
  1. 01The Lancet Gastroenterology & Hepatology — Routine AI assistance may lead to loss of skills in health professionals who perform colonoscopieseurekalert.org
  2. 02Microsoft Research / Carnegie Mellon — The Impact of Generative AI on Critical Thinking (CHI 2025)microsoft.com
  3. 03METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivitymetr.org
  4. 04FAA SAFO 17007 — Manual Flight Operations Proficiencyfaa.gov
  5. 05OpenAI — Usage policiesopenai.com
  6. 06OpenAI — Introducing study modeopenai.com
  7. 07Google — Guided Learning in Geminiblog.google
  8. 08Anthropic — Claude plans and pricingclaude.com
next guide
Everything your coding agent reads is an instruction
9 min · verified 2026-09-05
related guides