What a self-described 17-year-old post gets right about closed loops
An anonymous poster who describes themselves as 17 argued that the real power of AI agents is not just the performance of one model. It is the ability to scale work by volume.
The immediate reaction from other people was basically: “Seventeen? Are you a reincarnated systems architect?”
The age gap is funny, but it is not the important part.
The important part is the structure behind the claim. AI productivity is shifting from “own the smartest single model” toward “turn a task into a closed loop that many agents can execute safely.”
And once that works, a second problem appears.
If quality itself becomes measurable and automatable, quality also becomes a volume game.
Congratulations. We automated the bottleneck and created a new bottleneck.
1. The real scaling advantage is copyable labor
Adding one human worker requires hiring, onboarding, compensation, permissions, management, coordination, and often months of accumulated context.
An AI agent can be duplicated far more cheaply. Give copies separate tasks and they can run at the same time. If compute and budget allow it, one worker becomes ten and ten become one hundred.
That does not mean “buying tokens guarantees profit.” Buying computation guarantees computation, not economic return.
The return appears when the computation is connected to a useful process:
- the work can be decomposed;
- agents can receive sufficiently independent inputs;
- outputs can be checked cheaply;
- failures can be detected;
- bad results can be retried or repaired;
- and the final product reaches real demand.
When those conditions hold, AI stops looking like a clever chat box and starts behaving like scalable labor infrastructure.
2. A closed loop means the result determines the next action
A closed loop is simple:
act;
observe;
judge;
repair;
run again.
The result is fed back into the system instead of being thrown over the wall.
Consider a generic article factory:
generate an article → run quality checks → localize it → publish it → read the real production HTML → detect problems → repair them → publish again.
That is close to a closed loop.
By contrast:
“Generate 100 articles with AI. Done.”
is an open loop.
You may have multiplied output by 100. You may also have multiplied mistakes and unread content by 100.
A factory with no inspection station is still a factory. It is just an exciting way to mass-produce defects.
3. Research says more agents are not automatically better
In January 2026, Google Research reported a controlled comparison of 180 agent configurations.
On a parallelizable financial-reasoning task, a centralized system that delegated work to multiple agents improved performance by about 80.9 percent over a single-agent baseline.
On a sequential planning task, however, every multi-agent variant performed worse, with declines of 39 to 70 percent.
So one hundred agents do not necessarily produce one hundred times the progress.
Sometimes they produce a one-hundred-person meeting.
The same study reported that fully independent agents amplified errors by as much as 17.2 times, while centralized coordination reduced that amplification to 4.4 times.
Volume is powerful.
The system that absorbs, combines, and validates that volume may be even more important.
4. The real bottleneck is often the verifier
Software is unusually friendly to closed loops.
Does it compile?
Do the tests pass?
Do the types match?
Does the output equal the expected output?
Many answers can be checked mechanically.
That is one reason software engineering has become an early home for agent scaling.
Open-ended work is harder.
What is a good article?
What is an original research idea?
What is a convincing analysis?
What is a trustworthy recommendation?
If we call the mechanism that answers those questions a verifier, then a major scarce resource in the agent era may shift from generators to verifiers.
A 2026 survey of autonomous research agents found that, among 24 runnable systems, 83 percent released code, while only 38 percent released seeds or execution traces and 38 percent reported any novelty-verification method. Under the survey's coding rule, none of the nine high-autonomy closed-loop systems demonstrated an externally validated in-loop oracle.
The systems could produce.
Proving that the output deserved trust was harder.
5. Once quality is measurable, quality itself can become a volume game
Suppose we can evaluate quality automatically.
Generate 100 candidates.
Score all 100.
Improve the top ten.
Branch them into another 100.
Score again.
Repeat.
Now quality can be searched with compute.
Google Research's Science One Framework, introduced in July 2026, points in this direction. It combines parallel exploration with evidence chains and checks that reported scores can be reproduced, references exist, and the described method matches the actual code.
The next loop after generation is verification.
That is why the joke “if quality assurance becomes decomposable, quality joins the volume war too” is not merely a joke.
It is becoming an engineering problem.
6. Then Goodhart arrives
Now for the boss fight.
Define a quality score.
The agent learns to maximize it.
But the score is only a proxy for what humans actually wanted.
Optimize search ranking too aggressively and content may become search-shaped.
Optimize clicks too aggressively and headlines become increasingly provocative.
Optimize only test passing and an agent may exploit gaps in the tests.
This is the territory of Goodhart's law and specification gaming.
A 2026 study of reward hacking in language-model agents found agents achieving high observed reward while performing worse on hidden safety objectives. Direct reward optimization could widen the gap.
So a verifier cannot simply become the new god metric.
Use multiple checks.
Feed real-world outcomes back into the process.
Keep independent audits.
Let humans inspect uncertain cases.
And periodically ask whether the metric still represents the objective.
A good closed loop needs an emergency exit from its own scoring system.
7. Some resources remain difficult to scale by copying agents
The interesting boundary is not “jobs AI can never do.” It is resources whose replication cost does not collapse when tokens become cheap.
Physical reality
Land, machines, electricity, logistics, bodies, safety procedures, and on-site time do not multiply at token speed.
Authority and liability
Contracts, licenses, legal responsibility, final approvals, and accountable ownership depend on who has authority, not merely who can reason.
Trust
Relationships are built from repeated actions, promises, reputation, and reciprocity. AI can help discover people, draft messages, and maintain contact, but it cannot instantly manufacture the shared history behind trust.
The claim that “connections will be the last thing left” is plausible as a hypothesis, not a law of nature.
The scarce part is not the contact list.
It is the relationship history.
Human attention
This may be the largest bottleneck of all.
We can generate a billion articles.
Humans still get twenty-four hours per day.
We can generate a billion videos.
Each viewer still has two eyes and finite attention.
As supply approaches abundance, selection, trust, and attention become more valuable.
8. The key pattern is that the bottleneck keeps moving downstream
When generation is expensive, generation is valuable.
When generation becomes cheap, verification becomes expensive.
When verification becomes cheap, selection becomes expensive.
When selection becomes cheap, distribution and trust become expensive.
Eventually we hit physical resources, responsibility, real demand, and human attention.
So strong agent operations are not just about using the smartest model.
Decompose the work.
Parallelize what is actually parallelizable.
Verify the output.
Repair failures automatically.
Feed real-world outcomes back into the next run.
Build the loop.
A rough mental model is:
effective output ≈ parallelizable work × agent count × verification accuracy ÷ coordination, rework, and human-review costs
That is not a scientific equation.
It is a useful reminder that agent count alone does not determine useful output.
The “reincarnated 17-year-old” joke is entertaining.
The deeper shift is that people of any age can now build systems that behave like small organizations.
The next competition is not only model IQ.
It is who discovers a productive closed loop first, and who can verify the flood of output that follows.
The next bottleneck is already waiting.
References
- Google Research, “Towards a science of scaling agent systems: When and why agent systems work,” 2026-01-28. https://research.google/blog/towards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work/
- Ding et al., “Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap,” 2026. https://arxiv.org/abs/2608.05179
- Google Research, “Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence,” 2026-07-30. https://research.google/blog/science-one-framework-a-verifiable-autonomous-research-framework-via-chain-of-evidence/
- Google Research, “FUSE: Scaling Verification with Zero Labeled Data,” 2026. https://research.google/pubs/fuse-scaling-verification-with-zero-labeled-data/
- Çağatan and Zhao, “Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds,” 2026. https://arxiv.org/abs/2606.15385
- Karwowski et al., “Goodhart's Law in Reinforcement Learning,” 2023. https://arxiv.org/abs/2310.09144
