I handed senior-level engineering to an AI agent from my phone—and the move finished first

Share this article
I handed senior-level engineering to an AI agent from my phone—and the move finished first
AI-generated image
Advertisement
Advertisement

A fairly large software change was delegated to an AI agent.

It looked like the kind of thing that might be done in a few minutes. Instead, GitHub kept accumulating commits, tests, integration changes, and follow-up fixes. Hours later, real code was still changing.

Meanwhile, the human finished other real-world tasks—eventually an entire move.

The life event completed before the refactor did.

That is already a strange picture of the future.

The stranger part is the interface: a smartphone. The person giving the instructions does not need to personally understand every TypeScript module, CI/CD detail, or graph-learning rule. They can specify the objective, invariants, permissions, forbidden changes, and completion criteria in ordinary language, while the agent reads the repository, designs changes, writes code, adds tests, opens pull requests, and integrates work.

So is this “becoming a senior engineer with only a phone”?

Not exactly. But it is closer than the old mental model of software development suggests.

1. The phone is not doing senior-level computation

The phone itself is not generating hundreds of lines of production code locally. It is the command console.

Behind it are remote AI models, GitHub, CI systems, cloud infrastructure, production environments, search, and development tools. The phone is simply the interface through which intent and constraints are transmitted.

A more accurate description is not:

“software engineering on a phone,”

but:

“orchestrating remote cognitive and computational resources from a phone.”

The data center did not shrink into a pocket.

The pocket gained a remote control for the data center and the agent.

2. Why this kind of work is senior-level

Difficulty is not measured by line count alone.

A change of this kind requires someone—or something—to:

  • understand an existing architecture;
  • avoid duplicating existing mechanisms;
  • preserve production paths;
  • understand boundaries among schedulers, GitHub, CI, and publishing;
  • connect new decision logic into an existing central loop;
  • protect privacy;
  • avoid interpreting missing data as zero;
  • add tests;
  • distinguish code failures from CI infrastructure failures;
  • decide not only what to change, but what must remain untouched.

That is not merely implementing a small ticket. It is managing the blast radius of change inside a living system.

For a human team, this leans toward a senior backend engineer or platform engineer. If the person owns system-wide architectural responsibility, some of the judgment approaches staff-engineer territory.

The hard part is not just writing clever code.

It is knowing where clever code is allowed to exist without causing an accident.

3. How large was the actual change?

In one anonymized implementation example, the main pull request included:

  • 11 changed files;
  • 13 commits;
  • roughly 895 added lines;
  • an autonomous germination runtime;
  • Article DNA preservation;
  • behavioral-learning logic;
  • bounded feedback into internal-link optimization;
  • tests;
  • a dedicated CI definition.

And the work continued after merge. Additional gates were added so behavioral learning could not rewrite production graph behavior before verification, and missing telemetry remained UNKNOWN rather than being silently interpreted as zero.

So this was not “an 895-line job.”

It was a job that required understanding the system before adding 895 lines, then building a cage so those 895 lines could not run wild afterward.

Line count is a poor proxy for software difficulty. A hundred lines can delete a database. Ten thousand lines can implement an enthusiastic calculator.

4. How long would a human take?

For a strong engineer unfamiliar with the repository, a rough estimate for the same responsibility scope would be around 5–15 person-days.

The work includes:

  1. reading the existing code and operating rules;
  2. designing the change;
  3. implementing it;
  4. testing it;
  5. diagnosing CI and infrastructure failures;
  6. responding to review;
  7. checking production impact.

Typing the code might be faster.

In production systems, however, proving that you did not break anything can cost more than writing the change.

Japan’s IPA explains a conventional conversion of one person-month as 160 person-hours—8 hours × 20 working days—when no other definition is specified.[1]

On that scale, 5–15 person-days is about 0.25–0.75 person-month.

5. What would it cost to hire a human?

There is no universal price. Contract type, responsibility, prior system knowledge, production warranty, and review expectations all matter.

But public market rates give a sense of scale.

Levtech states that hiring a freelance IT consultant full-time for five days a week was roughly ¥1.0–1.1 million per month, based on July 2025 data.[2]

Using 20 working days as a simple divisor gives roughly ¥50,000–55,000 per day. Applied to 5–15 person-days, direct labor alone lands around ¥250,000–825,000.

A real contracted project can be higher once project management, review, rework risk, guarantees, overhead, and margin are included.

So this is not reasonably described as “a tiny task worth a few thousand yen.”

But it would also be misleading to say “the AI earned ¥800,000.” AI has different speed, parallelism, failure modes, supervision costs, and tool costs.

The more useful conclusion is this:

work that could consume several days or weeks of highly skilled human engineering can now be initiated by one person from a very small device.

6. Is six hours of agent time equal to six hours of senior-engineer time?

No.

An AI agent does not take coffee breaks, get pulled into meetings, lose half an hour to chat notifications, or stare at the ceiling wondering who approved the architecture.

It can also move rapidly across tools.

But it can:

  • accelerate in the wrong direction;
  • misunderstand production state;
  • confuse CI infrastructure failure with code failure;
  • expand permissions beyond what was intended;
  • confuse “a test exists” with “the test actually passed.”

Therefore, agent time should not be valued by the clock alone.

Judge the result by the quality of the completed change, evidence, verification, and production readback.

“Six hours and still working—good persistence” is fine.

“Six hours means six hours of correct engineering” is not.

7. What happens if the user cannot code independently?

The role changes rather than disappearing.

Historically, someone with an idea often had to learn Git, a language, a framework, deployment, and testing before the idea could become software.

AI agents reduce that implementation friction.

That shifts the human’s high-value work toward questions such as:

  • What are we building?
  • Why does it matter?
  • What must never break?
  • What is the agent allowed to change automatically?
  • What counts as success?
  • What remains UNKNOWN?
  • When does one failure stop the system, and when should it be isolated?

“Can personally write every line” is becoming less exclusive as an entry ticket.

That does not mean technical understanding is irrelevant. Even if you cannot write every module yourself, understanding objectives, risks, dependencies, and verification makes your instructions much stronger.

You do not need to manufacture an engine to drive a car.

You still need to know what a red light means.

8. Why “just let AI handle it” is dangerous

The scariest failure is not an obvious error screen.

It is being wrong with the appearance of success.

Examples include:

  • a PR is merged but nothing reached production;
  • a CI file exists but no runner ever started a step;
  • unavailable data is treated as zero;
  • new logic is added while old logic still runs, causing double execution;
  • “safety” disables the automation entirely;
  • “autonomy” quietly expands privileges too far.

Good automation is not merely automation that keeps moving.

It is automation that distinguishes facts from unverified assumptions while moving.

That is a major dividing line between senior-grade operations and enthusiastic automation theater.

9. How to delegate safely from a phone

You do not need to memorize every framework. A better pattern is:

Define the outcome first

Do not only say “edit this file.” State what must be true when the work is finished.

State invariants

Write down what must not break: existing features, privacy, image rules, publishing routes, SEO, production conditions, and so on.

Define authority

Specify whether overwriting, opening PRs, merging, or production changes are allowed.

Make the agent read current state first

Prefer current main, real runtime, and real production state over old conversation summaries.

Make verification part of the deliverable

“Code written” is not completion. Tests, CI, and production readback may be required.

Allow UNKNOWN

Do not force missing information into PASS or FAIL.

This turns a phone prompt from a casual request into something closer to an operating contract.

The screen is small.

The responsibility of the specification is not.

10. Conclusion: maybe the entry barrier is what broke

A person who cannot independently write advanced production code can now, under the right conditions, direct an AI agent from a phone for hours and integrate changes containing senior-level architectural judgment into GitHub.

A few years ago, that sentence would have sounded absurd.

Now it can be operationally real.

That does not mean engineers are obsolete.

A better interpretation is that part of engineering value is moving from manual implementation toward architecture, constraints, verification, and responsibility boundaries.

AI does not make value disappear.

It changes where value sits.

And perhaps the biggest shift is that people who used to stop at “I cannot technically build this” can now begin one step earlier:

“Then what should be built?”

In a world where the move can finish while the agent is still coding, development is no longer identical to sitting at a keyboard and typing every character yourself.

A smartphone is still just a slab of glass.

But that slab can now act as a remote control for senior-grade engineering resources.

That is, admittedly, a little broken—in an interesting way.


Sources

  1. IPA, FAQ for the Software Development Data White Paper series. It explains the default conversion of one person-month as 160 person-hours (8 hours × 20 days) ipa.go.jp
  2. Levtech, “ITコンサルタントに業務を依頼した場合の費用相場は?選び方やコスト削減法,” updated 2026-08-18. It reports roughly ¥1.0–1.1 million/month for full-time freelance IT consultants based on July 2025 data levtech.jp
Advertisement

Find other articles

All articles

Mendoi-chan

Written by

Mendoi-chan

She turns friction at work and in everyday life into clear structure and practical next steps.

About
Advertisement

Latest articles

  1. 1Was an 18-Hour Sleep Day Recovery Sleep? How to Understand Long Sleep and Changes in Dreams
  2. 2Do You Really Need to Apologize for Not Giving Your Parents Grandchildren? Sometimes an Adult Child Coming Home for Dinner Is Already a Big Deal
  3. 3The Day a 40-Year-Old VTuber Became a “Digital Community Center”: Age Does Not Always Kill Demand—Sometimes It Changes Its Shape
  4. 4How AI Article Automation Turned Into an Autonomous Factory in About a Week: One Ultra Punch, Level 6, and Why Level 7 Can Wait
  5. 5AI Is Brilliant, but the Factory Stops at “So… What Are We Building?” — The Person Who Ignites the First Idea Turns Capability into Production

You may also like

Advertisement