The people in Anthropic's artifact sitemaps who never clicked publish

August 7, 2026 · Part B of B

The first post explained why ten thousand published Claude artifacts are sitting in Google’s sitemaps: they’re leftover SEO from an Artifacts Gallery that Anthropic built, populated, and then retired. The feature is gone. The sitemaps are still served. Selection was by popularity — nothing under 16 views made it in — and indexing is on by default, with nothing in the publish flow mentioning any of it.

This post is what that has cost.

I probed all ten thousand on 26 July 2026. A census, not a sample. The number that matters isn’t the ten thousand though. It’s that some of those documents belong to people who never touched Claude.

What’s in there

Searching titles across the census turned up 19 artifacts carrying real personal data, from roughly 17 identifiable people. All 19 are marked index, follow. Thirteen carry an email address, fifteen a telephone number, five a LinkedIn profile. Two carry a date of birth and a home address. They have between 46 and 355 views each.

Titles only find what announces itself. So I scanned full source on a seeded random sample of 748 live artifacts, using a detector I validated first: 16 of 19 known positives caught, zero false positives on 10 known negatives.

Twenty-six of 748 carry personal data. That’s 3.48%, 95% CI 2.38–5.04%, which across the live corpus implies roughly 390 artifacts. That figure is an extrapolation, not a count. Treat it as an order of magnitude. Remember the engagement floor from part one: this corpus was curated for popularity, so it under-represents the quiet personal document nobody looked at. The real number is higher.

Hand-classifying those 26 is where it stops being a statistics exercise:

Category Implied What it is
Third-party personal data ~75 Other people’s details, published by someone else
Self-published ~63 CVs, portfolios, a personal life-update newsletter
Business or organisational ~113 Landing pages, flyers, pitch decks — intentional
Incidental or unclear ~75 Academic author addresses, demo apps

The row that matters

The third-party category is the entire argument. What I observed directly, not inferred: a shopping-centre management contact directory. A list of fifty named individuals presented as candidates. A tenant report portal. A contact-management system. A community help app carrying requester details. A piece of correspondence.

Think about the fiftieth name on that candidate list.

They didn’t publish anything. They probably don’t know the document exists. No account, no notification, no setting to check. They can’t discover they’re indexed, because you can’t search for a thing you don’t know to look for. And there’s no route to removal. Unpublishing belongs to whoever clicked the button, and it’s irreversible, so even that person only gets one attempt at it.

Someone who published their own CV made a choice. Badly informed, maybe. Part one is about how little the product told them. But they clicked. The fiftieth name on the candidate list made no choice at all.

That’s the difference between an exposure and a decision. I think it’s why this category is the one to fix first, though I’d understand an argument that the self-published ones are more numerous and therefore matter more.

How I nearly missed all of it

My first pass through the resume-shaped artifacts concluded there were no filled-in personal CVs in the corpus — just builders, templates and optimisers. I wrote that down as a finding.

It was wrong, and the reason is worth more than the finding was. The CV tools are the popular artifacts: 25,132 views, 12,802, 8,559. Every real personal CV sits between 46 and 355.

Search ranking puts the tools on top and buries the people. Read the first page of results and you conclude there’s nothing there.

There’s a language version of the same trap. English-only CV vocabulary found 24 candidate titles. Adding Japanese, Korean, Indonesian, Chinese, Russian, Arabic, Hebrew and Thai found 39. Keyword choice determined the finding more than the corpus did — which is the single best argument for running a census instead of a sample.

There is already a control for this

What the artifact suppression gate withholds from Google, and what it does not Two columns comparing artifacts the per-artifact suppression control removes from search against artifacts it leaves indexed. The control fires on finance and gambling vocabulary at 29.9 percent versus 0.71 percent for everything else, so it withholds a Claude Code install guide and a child's revision plan while leaving personal CVs, contact directories and candidate lists in Google. A per-artifact switch decides what Google may keep. It reads vocabulary, not subject. Finance- or gambling-titled artifacts suppressed 29.9% 106 of 355 Everything else 0.71% 64 of 9,027 WITHHELD FROM GOOGLE — noindex, nofollow SUBMITTED TO GOOGLE — index, follow A Claude Code install guide for Windows 26,600 views · zero finance or gambling words in its source A school revision plan naming the student on the title's evidence, a school-age child — right outcome, wrong reason Motherboard flashcards · a vehicle maintenance scheduler a participant schedule · a neighbour-dispute illustration A children's sight-word game · a family trip planner fired on a spin mechanic, a bingo, a budget line, a split golf bill A named person's CV — email address and phone number confirmed present in Google's live index A list of fifty named individuals presented as candidates none of whom published anything A contact directory · a tenant report portal other people's details, published by someone else ~390 artifacts carrying personal data of which ~75 are third-party — measured, not estimated upward The control is not decorative. It removed a Windows install guide from Google. It left a named person's CV in it. What the 170 suppressed artifacts are titled for finance only — 82 gambling only — 42 neither — 43 The narrow bar is the 3 titled for both. Twelve of the 43 “neither” were fetched and scanned — six contain no finance or gambling vocabulary at all. Measured 26 July 2026 · census of all 10,000 sitemap artifacts, not a sample Odds ratio 59.6. Stale-snapshot and preview-image fields are identical across suppressed and indexed artifacts, so that explanation is excluded. Personal-data rate from a seeded random sample of 748 live artifacts, full source scanned, detector validated at 16/19 sensitivity and 0/10 false positives. No artifact in the corpus has fewer than 16 views, so every proportion here is a floor. bensimon .dev
A control that removes artifacts from search — and what it covers — click to enlarge

So this is the part I didn’t go looking for. It’s a separate system from the sitemaps.

Every artifact carries a per-artifact robots directive. For 170 of the live ones it’s set to noindex, nofollow — and that suppression is not random:

Suppressed Total Rate
Finance- or gambling-titled 106 355 29.9%
Everything else 64 9,027 0.71%

Odds ratio 59.6. I checked and excluded the boring explanation — stale-snapshot and preview-image fields are identical across suppressed and indexed artifacts.

It reads vocabulary, not subject matter. Of the 170, forty-three are titled for neither finance nor gambling. I fetched twelve of those and scanned their source: six contain literally zero finance or gambling vocabulary — a Claude Code install guide, motherboard flashcards, a participant schedule, a school revision plan, a vehicle maintenance scheduler, a neighbour-dispute illustration. The other six fired on incidental words: a Korean children’s sight-word game with a spin mechanic, a Japanese family-support service that runs a bingo, a donation page quoting amounts, a family trip planner with a budget line.

The revision plan deserves its own sentence. It names the student it belongs to, who on the title’s evidence is a school-age child. Withholding that from Google is the right outcome. But it wasn’t withheld for that reason. It was withheld on vocabulary, in the same batch as motherboard flashcards. Getting the right answer by accident isn’t the same as reasoning about the right thing.

Then I checked three artifacts against Google’s live index:

Artifact Directive In Google?
A clinical trial summary listing patients by name, age, ethnicity and skin type index, follow Yes, with a snippet
A personal CV — named individual, email and phone number in the document index, follow Yes
A Claude Code Windows install guide, 26,600 views noindex, nofollow No — zero results

I can’t establish whether those clinical records are real rather than demo data, and I’m not going to claim they are. The argument doesn’t need them. The CV I can. It’s a real person, their contact details are in Google, and they’re there because a default said so.

The third row is why this section is here. The control is not decorative — it demonstrably removes things from search. It just isn’t pointed at any of this. Whether that’s a decision or something nobody ever considered, I can’t tell you, and I’d rather ask the question than answer it for them.

What the sitemaps did well

I’d be misrepresenting this if I stopped at the harm.

Non-English artifacts are 36.8% of the live corpus. Spanish 819, Japanese 752, Korean 511, Chinese 272, Thai 245, Hebrew 213, and a long tail after that. Median views are flat across languages, between 77 and 88. Non-English artifacts get read about as much as English ones.

Now look at what tops those languages. The most-viewed Hebrew artifact in the entire corpus is a Claude Code installation guide for Windows 11. The most-viewed Korean one is a guide to agentic coding best practices at 11,090 views, followed by an MCP server guide and a Korean-language guide to Artifacts themselves.

In languages where Anthropic may not publish localised documentation, users wrote it themselves. A retired gallery’s sitemap is the reason anyone can find it. Across the corpus, 134 artifacts mention Claude, Anthropic, MCP or prompting: 1.4% of the artifacts and 27.7% of all the traffic.

So the abandoned machine has been doing something genuinely useful this whole time. Any fix that just switches indexing off breaks that, and I don’t think that’s the right answer either.

Five things I couldn’t answer

What I did about it

I sent all of this to Anthropic on 29 July 2026 — to their privacy address, copied to user safety — with artifacts identified by UUID and no personal data in the note itself. I gave them no deadline and asked for nothing in return. I held the referral list back rather than pushing an unsolicited spreadsheet of exposed people at them, and offered it on request.

Nine days later there’s been no response. I should be straight about that gap. Thirty days is the convention and I’m not meeting it. Part one went out on the 31st and this is the other half of the same work, so splitting them further to run a clock down felt like theatre. Reasonable people can think I should have waited. The disclosure stands open either way, and if Anthropic engages or ships a change I’ll say so here, dated, above this line.

I’m not contacting any of the individuals, and I want to be explicit about why. Their details came out of documents I was analysing. Using them to send unsolicited email seemed worse than not sending it. Notification, if it happens, is Anthropic’s to do. They have the account relationship and an in-product channel. The people in the third-party category can’t be reached by anyone except whoever published the document they appear in.

I haven’t published a single UUID or artifact URL, here or anywhere, and I won’t in the comments either. A list of links in one place converts scattered obscurity into a directory aimed at exactly the people this is about. Right now those documents sit behind random v4 UUIDs nobody can enumerate. That’s the only thing protecting the fiftieth name on the candidate list, and it isn’t much.

The fix I’d want isn’t complicated and it isn’t “turn indexing off.” Stop serving a cancelled feature’s sitemaps. Say what the publish button does, the way session sharing already warns you about credentials in a private repo. And point the control that already decides what Google may keep at the documents where somebody else’s name is the thing being published.

SOs

SOs — standing orders you can lift and prime your own agent with. (New here? The convention started with the Neverland post.) Paste them into a session for a one-time dose, or commit them to your CLAUDE.md / AGENTS.md and they stay in force.

``` STANDING ORDERS — Third-party data and honest sampling Source: “The people who never clicked publish” — bensimon.dev

  1. When a document contains a third party’s details, treat it as a different category from one containing the author’s own. The author made a choice. The third party did not, usually cannot discover the exposure, and has no route to removal. Say so before the document is published, shared, or indexed.

  2. Never characterize a corpus from the first page of search results. Ranking sorts by popularity, which systematically buries the long-tail items a survey is usually looking for. State the instrument: census or sample, the size, the seed, and the detector’s measured sensitivity.

  3. Separate counted findings from extrapolated ones every time you report both. Put the interval on the estimate. Never let an extrapolation travel in a headline, a title, or a summary line.

  4. Before any artifact ships, run a mechanical scan for names, email addresses and phone numbers that may have reached working files. Do not rely on judgment. Report what the scan found, including nothing.

  5. Do not publish identifiers — UUIDs, URLs, record keys — for anything sensitive, even when the point is that they are already public. A list in one place is a directory; scattered is not.

  6. When keyword search drives a finding, run it in the other languages the corpus contains before reporting. English-only vocabulary undercounted this survey by roughly 40%. ```