saturday, september 5, 2026 · the day's ai, attributed published by trilot llc · wyoming
guide · working with ai

What an AI citation is actually worth

Work out whether pay-per-citation licensing will ever pay a site your size, and which settings decide that long before any money is on the table.

Published 2026-09-05 · Updated 2026-09-05 · Read 9 min · Reviewed by Rami Steitieh

Verified 2026-09-04 · Rami
on this page · 0 / 0 checked

The question usually arrives after something small and irritating. A page you wrote gets summarised back at you inside an assistant, accurately, with your phrasing more or less intact and no link. Your referrals to that page were already drifting down. Then you read that Apple is negotiating nine-figure deals with news publishers [8] and that Microsoft has built a marketplace to pay publishers when AI systems use their work [4], and what you want to know is the boring part nobody writes about. Whether there is a queue. Whether it is worth standing in.

This guide is for a solo operator, freelancer or small team that publishes something an AI system might quote: articles, documentation, reference tables, product data, prices. It is not for a publisher with a rights department, it is not for anyone with a lawyer already on the file, and it is not contract advice. The useful shift is to stop reading these announcements as grants and start reading them as prices. A price only pays out in proportion to how often somebody buys, and the thing being bought here is not your archive. It is being the answer.

A flat licence and a per-use licence are bets on different things

On 12 August 2026, MacRumors reported on a Wall Street Journal story that Apple was discussing multiyear content agreements with news publishers to give the rebuilt Siri access to current information, with a nine-figure budget discussed for the payments [8]. The number is not the interesting part. The structure is. Apple is proposing agreements where publishers would receive payment when their content is used, which the report describes as different from standard AI deals [8]. The Siri in question is scheduled to arrive alongside iOS 27, iPadOS 27 and macOS 27 [8].

Set the news aside and keep the shape, because the shape outlives the deal. A flat annual licence pays you for existing inside a corpus. It is predictable on both sides, it is easy to budget, and it pays exactly the same whether your work turns up in a million answers or in none. A per-use licence pays you for being retrieved. It scales with your relevance to real questions, which sounds fairer and is a completely different risk. Under a flat fee, an obscure archive and a daily news wire are worth the same money. Under per-use, they are not.

If you are small, this distinction is not academic, because it decides whether you are in the market at all. Flat licences are negotiated one at a time by people with a rights department, and nobody is calling you. Per-use licensing is the only version of this that can be automated down to a site with one author, which is why the plumbing for it is being built at the standards and CDN layer rather than in conference rooms. That is the good news. The rest of this guide is the arithmetic.

The per-use plumbing exists, and most of it is not switched on for you yet

Two pieces of infrastructure matter. The first is a way to state your terms in a form a machine reads. Really Simple Licensing reached version 1.0 on 10 December 2025 with Recommendation status, which the specification defines as stable, broadly reviewed, and suitable for implementation and deployment [1]. It is an XML vocabulary you attach to a site through robots.txt, HTTP Link headers, HTML elements, an RSS module, or metadata embedded in media and data files [1]. The part worth reading closely is the <payment> element, whose type attribute can be purchase for a one-time payment, subscription for recurring access, training for payment every time the content is used for AI training, crawl for payment every time the content is crawled, use for payment each time the content contributes to an AI-generated output, contribution for good faith monetary or in-kind support, attribution where licensees must provide visible credit and a functional link rather than money, and free [1].

Read that list again and notice that pay-per-citation already has a name. It is use, and it sits next to crawl as a separate thing, because being fetched and being quoted are different events with different values. Deciding which of those you would rather be paid for is the whole negotiation, compressed into one attribute.

The second piece is a way to collect. Cloudflare opened pay per crawl in private beta on 1 July 2025, anchored on HTTP response code 402, which Cloudflare calls a mostly forgotten piece of the web [2]. There are two flows. In the reactive one, a crawler requests a page, gets a 402 Payment Required response with a crawler-price header, and retries with crawler-exact-price if it is willing to pay. In the proactive one, the crawler sends crawler-max-price upfront, and if the configured price is at or below that limit the content is served with an HTTP 200 OK and a crawler-charged header confirming the charge [2]. Publishers set a flat, per-request price across their entire site, and Cloudflare acts as merchant of record, aggregating the events and distributing the earnings [2].

Two caveats before you get comfortable. It is still in closed beta, and you have to ask to be let in [3]. And there is a trap in the documentation worth writing on your hand: if you block an AI crawler via Cloudflare’s WAF or Bot Management products, those rulesets override pay per crawl’s charge feature, and the blocked crawler gets no access to the zone at all [3]. Blocking and charging are separate decisions made in separate products, and the block wins. A site configured to block everything AI-shaped and also to charge for access is a site that will never be paid.

Three programmes that would pay you, and the unit each one counts

Microsoft announced the Publisher Content Marketplace on 3 February 2026, describing it as a solution that gives publishers a new revenue stream, provides AI systems with scaled access to premium content, and delivers better responses for consumers [4]. On payment, the language is that publishers will be paid on delivered value, with usage-based reporting that lets publishers understand how content has been valued [4]. Two details matter more than the framing. Microsoft states that the marketplace will support publishers of all sizes, from large national and international organisations to specialised and independent voices, and that participation is voluntary, with publisher-defined licensing terms [4]. It is also still a pilot, with demand partners beginning to be onboarded, and it has been co-designed with publishers including The Associated Press, Business Insider, Condé Nast, Hearst Magazines, People Inc., USA TODAY Co. and Vox Media [4].

Perplexity took a different route. Comet Plus is a $5 standalone subscription that also comes included with Pro and Max memberships, and Perplexity says it will distribute all of that revenue to participating publishers, minus a small portion for its compute costs [5]. The compensation categories are the informative bit: human visits, search citations, and agent actions [5]. That third one is new. It means a publisher can be paid because an agent read the page while doing something for a user, with no human ever landing on it.

Now line the three up. Cloudflare counts requests [2]. RSL lets you price a crawl event and an output contribution separately [1]. Perplexity counts visits, citations and agent actions as three different things [5]. Microsoft says delivered value and reports on usage [4]. Apple, per the reporting, wants to pay when content is used [8]. Every one of these is a per-use model, and no two of them are counting the same event. When you eventually read a contract, the unit is the deal. Everything else in it is decoration.

Run the number before you rearrange anything

Here is where most of these conversations should stop and almost never do. The click side of the ledger is already measurable, and it is not encouraging. Pew Research Center tracked the browsing of 900 US adults across 68,879 unique Google searches in March 2025 and found that users clicked a traditional search result on 8% of visits where an AI summary appeared, against 15% of visits where one did not [7]. Clicking a link inside the AI summary itself happened on 1% of the visits where a summary appeared [7].

So the honest framing is a swap, not a windfall. You are being asked to trade a referral that already converts poorly for a payment that scales with something you have never measured. Before you accept or refuse, measure it. Pull a month of server logs and count requests by user agent, separately from human sessions. That number, multiplied by a price you would actually charge, is the entire revenue line.

calculator
What a per-request price would pay you
$ / month

Requests × price ÷ 100. Cloudflare's pay per crawl applies one flat per-request price across your whole site [2], so this is the gross figure, before any crawler declines to pay it. Computed in the page; nothing is sent anywhere.

Run it with your real log numbers rather than the defaults. For most sites at this size the answer is a two-figure monthly sum, which is worth knowing precisely because it stops you making a large decision for a small amount of money. It also tells you which lever is worth pulling. If your crawl volume is high and your referrals are near zero, you are subsidising somebody and pricing is a reasonable response. If your crawl volume is modest, the licensing question is a distraction and your attention belongs elsewhere.

The crawler settings decide more than the cheque does

The instinct after reading all this is to block AI crawlers and wait for someone to offer money. That instinct is wrong in a specific, expensive way, because “AI crawler” is not one thing. OpenAI documents four, and they do different jobs [6]. GPTBot crawls content that may be used in training generative AI foundation models, and disallowing it signals that your content should not be used for that. OAI-SearchBot surfaces websites in ChatGPT’s search features, and sites opted out of it will not be shown in ChatGPT search answers, though they can still appear as navigational links. OAI-AdsBot validates the safety of pages submitted as ads on ChatGPT, and the data it collects is not used to train foundation models. ChatGPT-User covers certain user actions in ChatGPT and Custom GPTs, is not used to crawl the web automatically, and is not used to determine whether content may appear in search; because those actions are initiated by a user, robots.txt rules may not apply [6].

Each of those settings is independent of the others, so you can allow one while disallowing another [6]. This is the most useful sentence in this guide. The retrieval crawler and the training crawler are separate switches, and they buy you separate things. Blocking training is a rights position with no traffic consequence. Blocking retrieval removes you from the answers you were hoping to eventually be paid for appearing in. People conflate them constantly, publish one blanket rule, and then wonder why they stopped showing up.

The order of operations, then, is unglamorous and works. Set the training and retrieval switches deliberately and separately. Publish machine-readable terms even where nobody is paying yet, since RSL terms ride along in robots.txt and RSS and cost nothing to state [1]. Keep counting requests and referrals monthly, because that ratio is the only number you will bring to any negotiation you ever have. Then join the pilots that will take you, and treat whatever arrives as a rounding error until it demonstrably is not.

checklist
Before you sign anything or block anything
0 of 7 · saved in this browser only

What still goes wrong

Almost none of this is a product you can rely on yet. Cloudflare’s pay per crawl is in closed beta, and joining means going through a signup page or an account executive [3]. Microsoft’s marketplace is a pilot that is only beginning to onboard demand partners on the buying side [4]. Perplexity announced Comet Plus with three compensation categories and no published payout formula, saying only that revenue goes to participating publishers minus compute costs [5]. Apple’s arrangement is a reported negotiation, not a signed deal [8]. You cannot budget against any of it, and a plan that assumes this revenue arrives on a schedule is a plan with a hole in it.

A machine-readable licence is a notice, not a fence. RSL gives you a standard way to state terms, and the specification was edited by people from Condé Nast, Ziff Davis, Yahoo, Automattic, O’Reilly Media, Fastly and Schema.org, so it is not a fringe document [1]. Stating terms clearly is genuinely worth doing. But the file does not stop anybody. Enforcement lives at your CDN, and even there the controls fight each other, since a blocking rule silently beats a charging rule in the same account [3]. If you want the difference between a stated price and a collected one, that is it.

The last limit is the honest one about scale. Collective licensing works because it aggregates, and the arithmetic in the calculator above is the reason. A per-request price that produces meaningful revenue for a site with millions of pages produces a monthly sum you would not chase for an unpaid invoice at the size most readers of this guide operate at. The value of doing the work anyway is not the cheque. It is that you end up knowing your own crawl-to-referral ratio, which is the number that tells you whether AI systems are a distribution channel for your work or just a cost, and almost nobody running a small site can currently answer that.

sources
  1. 01RSL — Really Simple Licensing 1.0 Specificationrslstandard.org
  2. 02Cloudflare — Introducing pay per crawlblog.cloudflare.com
  3. 03Cloudflare Docs — What is pay per crawl?developers.cloudflare.com
  4. 04Microsoft Advertising — Building Toward a Sustainable Content Economy for the Agentic Webabout.ads.microsoft.com
  5. 05Perplexity — Introducing Comet Plusperplexity.ai
  6. 06OpenAI Docs — Bots (crawler user agents)developers.openai.com
  7. 07Pew Research Center — Google users are less likely to click on links when an AI summary appears in the resultspewresearch.org
  8. 08MacRumors — Apple in Talks to Pay Publishers for News Content to Power Siri AImacrumors.com
next guide
How to tell whether AI is actually working for you
9 min · verified 2026-09-04
related guides