Locale: en
On September 24, 2026, Japan’s National Tax Agency carried out a major renewal of its core tax systems. Soon afterward, tax-office counters began experiencing substantial delays in tasks such as receiving cash payments and issuing tax-payment certificates, while several e-Tax-related problems and emergency maintenance events appeared around the same cutover. Anyone who has built even a modest integrated system can sympathize with the engineering difficulty. But “a failure is understandable” and “the fallback was good enough” are two different judgments.
Imagine someone trying to complete a tax procedure quickly. The office says the broader system is having trouble, there is no firm recovery time yet, and an urgent case can still be handled in person.
At first, the answer sounds obvious: go to the office.
Then the queueing problem appears.
People who would normally finish online start going to the counter. At the same time, employees need more manual checking because the systems they rely on are degraded. More arrivals, less processing capacity.
“Come in and we can handle it” does not mean “come in and it will take five minutes.”
If the matter is not time-critical, waiting can be the rational choice.
And the natural reaction is still:
“Fine. Just use paper.”
That instinct is half right. Paper can be an escape hatch. It is not a backup database.
1. What happened in September 2026
The National Tax Agency renewed the national tax system on September 24, 2026. Its next-generation KSK2 program had been described around three major concepts:
- move administrative work from paper-centered processing to data-centered processing;
- integrate databases and applications that had been separated by tax category;
- move away from a proprietary mainframe-centered environment toward an open system using general-purpose operating systems.
This was not a cosmetic upgrade. It changed data organization, application boundaries, and core infrastructure at the same time.
On September 24, the NTA published a notice titled “Delays in various procedures at tax office counters.” It stated that cash receipt, issuance of tax-payment certificates, and other counter procedures were taking a considerable amount of time. Even certificates requested through e-Tax could not necessarily be issued immediately. The system renewal itself had been completed, but problems had occurred in systems required for counter operations.
During the same cutover period, there were problems with login via the Mynaportal route, an issue in the completion display for internet-banking tax payments, emergency maintenance, and a period when some e-Tax functions were unavailable. The e-Tax functional outage was reported resolved by September 27.
As of September 28, the counter-delay notice was still shown as emergency information on the official site. A detailed root-cause analysis had not yet been published.
It also matters to separate planned downtime from failures. The renewal already included a scheduled e-Tax shutdown from September 19 at 00:00 until September 24 at 08:30, plus a full-day shutdown on September 26. Post-cutover incidents and emergency maintenance were added on top.
So the feeling that “it has been down forever” is understandable, but the period combines both planned maintenance and unplanned disruption.
2. A serious money system can fail — and sometimes its seriousness is why it must stop
A common reaction is: “If this system handles taxes and money, shouldn’t it be built never to stop?”
Yes, high availability matters.
But critical systems also need integrity.
For a payment-related system, the dangerous outcomes are not limited to a page being unavailable for five minutes.
Much worse would be:
- a payment being made but recorded as unpaid;
- one operation being registered twice;
- a certificate being issued from stale data;
- one subsystem being updated while another remains behind;
- a recovery retry repeating an operation that already succeeded.
In those conditions, “keep processing somehow” can be more dangerous than stopping.
Google’s SRE guidance emphasizes that during a major incident, the first goal is to stop further damage rather than immediately solve the root cause. If continued operation risks unrecoverable data corruption, freezing the system can be safer.
The paradox is simple.
The system does not stop despite dealing with money.
Sometimes it stops because dealing with money makes an incorrect state unacceptable.
If “Have you tried restarting it?” solved everything, operating national infrastructure would be a much happier profession.
3. Every component can pass and the integration can still explode
Integration failures are unpleasant because individual components can all look healthy.
System A works.
System B works.
The database works.
Authentication works.
Yet if A hands B a slightly different format than expected, the workflow can fail.
A single exceptional record from a legacy migration can expose an unseen path.
A retry can lose the response even though the underlying operation succeeded, leaving the caller unsure whether to try again.
Old caches, historical forms, external institutions, permissions, batch jobs, character sets, rare names, timing windows, and replay behavior all create boundaries where defects can hide.
Even small personal automations run into stale state, duplicate jobs, dropped events, external API changes, and retry side effects.
Now scale that to a national tax system spanning tax categories, collection, refunds, certificates, e-Tax, offices, and external organizations.
People who have experienced integration pain tend to look at a large cutover and think, “Yes, something surfacing right after launch is believable.”
That is not an excuse.
The more useful question is not whether the migration produced zero defects. It is how safely the service degraded when defects appeared.
4. Why recovery can have no reliable ETA
Few phrases are more frustrating to users than:
“We do not have an estimated recovery time.”
But an invented deadline can be worse.
Recovery of a critical system is not finished when a server process starts again.
Operators may need to determine the affected scope, verify whether any writes were only partially committed, check whether retries can create duplicates, reconcile state with external systems, decide whether rollback is still safe, and then process the backlog accumulated during the outage.
One of the hardest states is: “The sender appears to have succeeded, but the receiver did not confirm final completion.”
That is why idempotency matters. Amazon’s Builders’ Library describes idempotent operations as a way to make retries safe: repeating the same intended request should not create additional side effects.
A missing ETA does not prove that nobody knows what they are doing.
If the team does not yet know the true blast radius, the integrity of in-flight transactions, or the safe replay point, a precise prediction may simply not exist.
5. “Come to the office if it is urgent” can create a brutal queue
In-person service is an important fallback.
But during a system incident, it can become a bottleneck very quickly.
Let the arrival rate of visitors be λ and the counter’s processing capacity be μ.
An incident can push both in the wrong direction at once.
People who would normally use online channels move to the office, increasing λ.
Employees need extra checks, manual work, and exception handling, reducing μ.
Demand rises while service capacity falls.
That is the worst possible direction for a queue.
And tax-office staff do not automatically multiply when a server goes down.
Google SRE guidance notes that overloaded systems often become nonlinearly worse near capacity: latency rises, dependencies time out, and cascading problems can appear. Graceful degradation and load shedding exist because “a little more load” can become much more than “a little slower.”
So “we can handle it at the counter” is not evidence that the counter is quiet.
If there is no hard deadline, avoiding the peak disruption window can save a large amount of time. If a legal deadline or genuine urgency is involved, however, users should not simply wait on assumption; they should follow the current official instructions for alternative procedures.
6. “Just use paper” is half right: paper can preserve intake, but it cannot become the database
When systems fail, the most intuitive reaction is:
“Take the form on paper.”
That is not a foolish idea.
NIST contingency-planning guidance explicitly recognizes manual processing as one possible short-term method for continuing business processes during an information-system disruption.
Manual processing is a real resilience tool.
But paper cannot necessarily finish the whole transaction.
Paper may be useful for:
- proving that a request was received;
- fixing the receipt time;
- collecting required documents;
- establishing an order for later processing;
- issuing a receipt or reference number.
What it may not be able to do alone is confirm the central tax record, verify payment state, consult historical data, issue an accurate certificate, or complete an external-system exchange.
The ideal therefore is not “run the entire tax system on paper.”
It is “do not let intake die just because the core system died.”
Paper is not a backup database.
But it can be an emergency exit.
7. What users actually need is graceful degradation
A resilient service does not always preserve 100 percent of normal functionality.
Sometimes it deliberately falls back to a smaller, safer mode.
Google SRE calls this graceful degradation. NIST contingency planning similarly includes alternate equipment, alternate locations, and manual processing among possible recovery strategies.
For a tax procedure, a strong degraded mode would look like this:
- Keep intake alive. Accept the minimum required request data offline or on paper.
- Issue a receipt identifier. Remove uncertainty about whether the office received the request.
- Store work durably for later processing. Recovery should not depend on someone remembering a pile on a desk.
- Make replay safe. The same request should not create duplicate legal or financial effects.
- Prioritize urgent cases. Deadlines and serious personal consequences should be distinguishable from ordinary work.
- Publish precise status. Separate what is down, what still works, and what has recovered.
- Reconcile after recovery. Compare offline/manual intake with the restored production records and detect omissions or duplicates.
The engineering goal is not the fantasy that nothing ever breaks.
It is to design how the system breaks.
8. How often does e-Tax actually have problems?
The official e-Tax notice list for 2026 shows multiple incident notices across the year: delayed payment-completion notifications in January, Mynaportal-linkage and login problems in February, login difficulty in March, a Mynaportal-related problem in July, direct-payment trouble in August, and several incidents around the September system renewal.
But these should not be thrown into one bucket.
Some involve e-Tax itself. Some involve Mynaportal integration. Some involve payment functions or display behavior. Planned maintenance is not an outage.
The defensible conclusion is therefore narrower: partial failures and integration problems are publicly reported multiple times in a year, but that does not prove that the entire national tax system is constantly suffering nationwide total outages.
September 2026 stands out because a very large cutover, long planned maintenance windows, and multiple post-cutover problems occurred in the same week.
9. Once you have built integrations yourself, the way you get angry changes
Anyone who has built even a small automation stack tends to react differently to big outages.
Before that experience, the response may be:
“How can this possibly be broken?”
After connecting several services, queues, retries, persistent state, and external APIs, another thought appears:
“Oh. A major integration cutover. That sounds painful.”
Everything works alone, but the full flow fails.
Fix one boundary and another appears.
Old state survives.
A retry becomes a duplicate.
The log you trusted turns out to describe yesterday.
Experiencing the miniature version makes the difficulty of national infrastructure easier to understand.
But understanding is not immunity from evaluation.
The right questions remain:
- How thoroughly was migration and overload behavior tested?
- How far could the service degrade safely?
- Did manual intake actually work?
- Was public status communication specific enough?
- Will a root-cause analysis and corrective action follow?
- Will the same failure class be removed before the next major cutover?
Complexity makes incidents plausible.
Complexity also makes learning from them mandatory.
10. Conclusion: not “run everything on paper,” but “do not let the front door die”
When a national tax system stalls, users experience a particularly annoying contradiction.
This is a place that handles money and legally important procedures, yet the system can still stop.
There may be no recovery estimate.
The office may say in-person handling is possible, while every instinct says the counter will be packed.
That is why “just use paper” feels so satisfying.
The useful requirement hidden inside that joke is not that paper should replace the whole tax platform.
It is that an outage in the core system should not automatically destroy intake, evidence, prioritization, and the ability to replay work later.
Large systems do not need a myth of perfect invulnerability.
They need a safe failure mode.
They need to recover without duplicating transactions.
They need to tell users exactly what works and what does not.
And after recovery, they need to process everything that accumulated while the system was down.
So the final demand is not:
“Do the entire tax system on paper.”
It is:
“At least let the office receive the request on paper. Seriously.”
Sources
- National Tax Agency, “Delays in various procedures at tax office counters,” 2026-09-24
https://www.nta.go.jp/files/000041014.pdf - National Tax Agency, “Renewal of the national tax system”
https://www.nta.go.jp/taxes/shiraberu/sodan/system.htm - National Tax Agency Report 2025, KSK2 overview
https://www.nta.go.jp/about/introduction/torikumi/report/2025/03_5.htm - e-Tax, maintenance associated with the national tax-system renewal
https://www.e-tax.nta.go.jp/topics/2026/topics_20260422.htm - e-Tax notices
https://www.e-tax.nta.go.jp/topics/ - e-Tax, resolved service unavailability notice
https://www.e-tax.nta.go.jp/topics/2026/topics_20260925_mentenansu.htm - NIST SP 800-34 Rev.1, Contingency Planning Guide for Federal Information Systems
https://csrc.nist.gov/pubs/sp/800/34/r1/upd1/final - Google Site Reliability Engineering, Handling Overload / Effective Troubleshooting
https://sre.google/sre-book/handling-overload/
https://sre.google/sre-book/effective-troubleshooting/ - Amazon Builders’ Library, Making retries safe with idempotent APIs
https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/
