I write articles with AI. I translate with AI. I fix broken links with AI. I tally traffic with AI. I ask AI, "Is this part confusing?" and then fix it with AI again.
Get far enough down that road and you hit the same problem.
Everyone involved knows the site too well.
The people who built it know what every button does. The AI reads the spec and understands that "this button is for related articles." The site owner has looked at the pages so many times that their brain quietly fills in anything slightly off.
A first-time visitor doesn't have any of that.
"What is this?" "Where do I click?" "Why is only this part in English?" "I can't go back." "I can read the text, but something about it feels off."
That "something" is exactly what I want.
So I came up with an idea, and it's a pretty funny one.
I'm finally outsourcing the debugging step to actual humans.
And I'm paying 1 yen per finding, with a cap of 1,000 yen per person.
Humans are coming for AI's jobs. Their department: "a gut feeling that something's wrong."
1. What I want to outsource isn't a code review, it's "the first-time snag"
I'm not looking for a big expert audit here. I want things like this:
- I can't tell what a button means
- I don't know where to go next
- A heading is oddly worded
- It's hard to tap on a phone
- The link goes to a page in a different language
- I tapped something and the screen went blank
- The sentence is grammatical, but a human reading it thinks it's strange
- It'd be handy to have a feature here
Digital.gov describes usability testing as watching real users try to use a product or service. [1] GOV.UK also says that watching real users work with a service helps you find concrete problems, such as wording and layout. [2]
So "let a human touch it once" isn't some strange ritual that suddenly appeared in the AI era.
It's the old usability test, coming back to a site that AI can now churn out and patch in bulk.
2. Sometimes the site owner is the worst person to test it
Check your own site carefully every day.
Easy to say. Once articles, languages and features keep piling up, it's simply not possible.
And "the person who built it does the checking" has its own problem.
They know where the search box is. They know where related articles show up. They know how language switching is supposed to work. Even when something's a bit off, they wave it through as "that's just how it works."
And above all, looking is tedious.
That matters more than it sounds.
If you try to beat the tedium into submission with willpower, as in "let's check 100 pages every day," that routine won't last. So instead, split the tedious part off from the rest.
The owner doesn't need to be the person who looks at everything.
They can be the person who decides what to look at and builds the system that fixes whatever turns up.
3. What 1 yen per finding buys isn't ideas, it's fresh eyes
On its own, 1 yen per finding sounds absurdly cheap.
But I'm not trying to buy a polished consulting deck.
I'm buying the fact that a fresh pair of eyes got stuck once.
If one person reports 1,000 things, that's 1,000 yen. But rewording the same finding still counts as one. And the payout per person is capped at 1,000 yen.
That cap matters a lot.
The moment you say "I want lots of reports," the world immediately starts heading toward "so if I make AI spit out 10,000 improvement points, that's 10,000 yen, right?"
So alongside the per-finding payment, I set these conditions:
- You actually looked at the site
- You write about real places and real behavior
- You don't pad the list with rewordings of the same point
- There's a cap
One more note: 1 yen per finding is just the design of this particular experiment, not a universal fair rate. If you want long hours of work or expert evaluation, the terms should match the time and difficulty involved.
4. The biggest enemy is "I had AI come up with 10,000 improvements"
There's no need to ban AI.
The problem is improvement ideas from someone who never looked at the site.
"In general, websites should make navigation clearer." "Let's improve accessibility." "Let's optimize for responsive layouts."
All true.
And none of it is what I'm after.
What I want is observations like:
"I couldn't tell what would happen if I pressed this button on this page." "Right after this heading, the topic suddenly jumped." "Only in this language version, the related articles link to a different language."
AI is useful for trimming those observations down and merging duplicates.
Humans discover, AI organizes.
What matters is not flipping that order.
5. If reports arrive one at a time in chat, the requester dies first
If you want more reports, you have to design how they're submitted, too.
The worst setup is this:
"Here's number one." "Number two." "Oh, and number three." "One more, number four."
The chat just keeps growing forever.
You outsourced the debugging, and now you've created a new manual job: collecting the reports.
So submissions come in as one batch. Any of these works:
- Excel
- A spreadsheet
- A .txt file
- A numbered list
- Text that can be copy-pasted
Screenshots are basically unnecessary. They're handy to look at, but they cause a lot of friction later when you want to search, remove duplicates or run AI over them.
GOV.UK's user research guidance also says typed notes are easier to store and analyze afterward, and that splitting each observation into its own item makes sorting and analysis easier. [3]
The format can be simple.
Where / what it looks like now / what would make it better
For a bug:
Where / what you did / what happened
That alone gives the AI downstream plenty to work with.
6. On a multilingual site, "it got translated" and "a human can read it" are different things
Multilingual sites get even more interesting.
By machine standards, all 12 languages were generated successfully. The build succeeded. HTTP 200. The links exist.
And yet, a human looking at it can still find plenty that's off.
- Some of the original language is left behind
- Only the buttons are translated oddly
- The sentence makes sense, but sounds unnatural to a native speaker
- Longer text breaks the layout
- After switching languages, only the related articles go back to the original language
- It's supposed to be the same article, but parts are missing
Only people who can read the language need to check it.
Those who read English check English. Those who read Korean check Korean. Same for Chinese, Spanish, Portuguese, Indonesian, Thai, Vietnamese, French and German.
Nobody has to look at every language.
The machine confirms "it was generated," and a human confirms "does this read normally?"
They do different jobs.
7. A duplicate counts as one for payment, but it matters enormously for analysis
Say the same person writes:
"The button is confusing." "I don't understand what the button means." "What is this button?"
That's one finding.
But if three separate people each get stuck on the same button on their own, that's a different story.
That's not just a duplicate. It's friction that can be reproduced.
So when tallying, I split it like this:
- Rewordings from the same person: merge into one
- The same finding from independent people: keep it as a count
"Three people got lost in the same spot" carries more weight than the owner's personal taste.
GOV.UK's Service Standard also calls for checking usability often with real or likely users, and testing on a range of devices that reflects how users actually behave. [4]
The point of having more people isn't to hold a vote on opinions. It's to see whether the same friction happens to other people too.
8. The finished loop is AI, then human, then AI, with the human as the sensor
The final process turns out pretty clean.
- AI generates the article
- AI translates it
- AI checks the links, structure and display
- AI fixes the known problems
- A human actually reads, clicks and gets lost
- The human writes down "something feels off" as text
- AI merges the duplicates
- AI sorts them by severity, reproducibility and cost to fix
- AI drafts candidate fixes
- After the fixes, it goes back to machine checks and human review
At last, even the human has become one step in the pipeline.
But the work left to humans isn't mindless labor.
It's detecting the friction that only a human can feel.
If anything, it's closer to a promotion.
Let the machine check "is the link a 404?" and let the human check "it's not a 404, but I don't understand why I was sent here."
Let the machine check "does the translated string exist?" and let the human check "it exists, but nobody actually says it that way."
That's the more natural way to split the work.
9. Testers viewing pages also raises page views, but that isn't the goal
Naturally, if people go through dozens of pages, the view count goes up.
On a site with ads, ads may also be displayed as part of normal browsing.
But that's purely a side effect.
Asking people to click ads, or making more ad impressions a condition of the work, would be a different matter entirely.
What I'm buying isn't page views.
It's the observations left behind after a human looked at the page.
If traffic picks up a bit along the way, it's best to think of it as "I was doing user testing and the cash register rang a little, too."
10. The endpoint of automation turned out not to be "zero humans"
Once you start using AI, it's tempting to wonder how far you can erase humans.
But when you actually push automation forward, you sometimes land on the opposite conclusion.
You don't need to remove every human.
You just need to move humans to the places where only humans add value.
Drafting articles. Translation. Merging duplicates. Checking links. Sorting. Tallying. Proposing fixes.
Push all of that toward the machines.
And at the very end, buy only these from humans:
"First time seeing this, and I have no idea what it means." "I hate this part." "Something's off." "This is handy."
What 1 yen per finding buys isn't a line of text.
It's another person's split-second feeling that something's wrong, which neither I nor the AI can produce.
Push automation to the limit, and humans came back.
And their department is "something feels off."
Very human indeed.
References (4)
- Digital.gov, “Usability testing digital.gov
- GOV.UK Service Manual, “Using moderated usability testing gov.uk
- GOV.UK Service Manual, “Taking notes and recording user research sessions gov.uk
- GOV.UK Service Manual, “4. Make the service simple to use gov.uk



