Load Testing a Small Web App: DDoS, Concurrency, Recovery, and AI Coding Costs

AI can turn an idea into a working multiplayer app absurdly fast. Rooms open. Users join. Posts appear. Votes count. The brain immediately says, “Done.”

How reading tools work

Listen reads the article aloud. Speed read shows phrases in sequence at your chosen pace. Language practice compares available translations. Save keeps a bookmark in this browser; find it in the player’s bookmarks.

Share this article
Advertisement
Advertisement

The dangerous moment is not when the app fails. It is when one successful demo makes you think it is finished.

AI can turn an idea into a working multiplayer app absurdly fast. Rooms open. Users join. Posts appear. Votes count. The brain immediately says, “Done.”

The server says, “That was one user.”

A real pre-launch review has to ask different questions: What happens when 100 people act at once? What if a write succeeds but the response is lost? What if the deadline and the last vote happen together? What if a payment webhook arrives twice? What if everyone reconnects at the same moment after an outage?

A code review of one small multiplayer web app expanded into 324 test cases. That did not mean 324 bugs had been found. Every item was still unexecuted. It was a test plan, not a pass certificate.

A beautiful bridge drawing is useful. It is still not the same thing as driving trucks over the bridge.

1. DDoS is not the only way a small app gets crushed

Cloudflare provides automatic DDoS detection and mitigation on all plans, but Cloudflare also recommends application-level controls such as rate limiting because large attacks can still affect an application.[1]

For a small app, legitimate traffic may become the first real load problem.

If a client polls every three seconds, the raw request count looks like this:

Concurrent clients State checks per hour
10 12,000
30 36,000
100 120,000
1,000 1,200,000

This is not a capacity forecast. Backoff and caching can reduce it; posting, images, inbox requests, multiple tabs, and scheduled jobs can increase it.

As of October 5, 2026, Workers Free lists 100,000 requests per day. D1 Free lists 5 million rows read per day, 100,000 rows written per day, 500 MB per database, and 50 queries per Worker invocation.[2][3] A single D1 database processes queries one at a time; excessive concurrency queues and can eventually produce overloaded errors.[3]

So the scary traffic pattern is not only “a million attackers.”

Sometimes it is 100 perfectly legitimate users who are all having fun at the same time.

2. Why the checklist reached 324 items

The checklist covered sixteen areas: entry points and abuse, capacity, concurrency and recovery, scheduled jobs, rooms and permissions, images, text input, voting, deadlines, result release, public-room flow, persistent connections, devices and reconnection, external integrations, payments, and deployment monitoring.

The point is not to test every item with equal urgency. The useful move is to identify the failure modes that can corrupt data, leak permissions, stop game progress, burn quota, or charge money incorrectly.

3. The 23 checks worth attacking first

  1. Can rejected traffic still trigger expensive database work?
  2. Does your own usage counter actually count every expensive route?
  3. Does rate limiting survive restarts and multiple execution locations?
  4. Does “nothing changed” polling still read lots of database state?
  5. Can one seat open unlimited persistent connections?
  6. Do HTTP and persistent-connection actions enforce the same limits?
  7. Does an emergency stop affect clients that are already connected?
  8. Can a late action slip through because scheduled deadline processing was delayed?
  9. Can one failed side effect be silently skipped while the room advances?
  10. Can repair jobs double-count statistics?
  11. Can a start reservation be consumed before the actual game is created?
  12. Can a retry turn one answer into two?
  13. Can visually equivalent usernames bypass uniqueness rules?
  14. Can two users take the final available seat simultaneously?
  15. Can one scheduled job accumulate too much work to ever catch up?
  16. Are database size and indexes measured, not just uploaded images?
  17. Can somebody else take over an abandoned device identity?
  18. Are image type, dimensions, metadata, and location data handled safely?
  19. Do long histories make each page load increasingly expensive?
  20. Can synchronized reconnects cause a second outage?
  21. Have duplicate, delayed, failed, cancelled, and restored payments been tested?
  22. Are local tests being mistaken for production guarantees?
  23. If the database fails, does the monitoring path fail with it?

That list is much more useful than “I clicked every button once.”

4. Bugs love combinations

A high-value test looks like this:

100 users vote together
→ 20 refresh
→ 10 lose connectivity
→ some writes succeed but their responses disappear
→ those users retry.

Now check whether vote counts duplicate, late votes slip in, screens roll back to stale state, or a successful operation is performed twice.

OWASP treats concurrent sessions, abnormal input, and error handling as independent security testing concerns for a reason.[4]

5. “The write succeeded” and “the user received success” are different events

A database write can succeed while the network response disappears.

From the user's perspective, the action failed. They tap again.

Without an operation identifier or equivalent duplicate protection, the server may politely create a second answer, second room, second vote, second notification, or second charge.

Retries must be designed, not merely tolerated.

6. Load testing should be staged

Use an environment you control. Increase load in steps such as 10 → 30 → 100 clients.

Then vary the shape of those 100 clients:

  • one crowded room versus many rooms;
  • simultaneous joins;
  • simultaneous final votes;
  • synchronized reconnects;
  • scheduled jobs at the same time;
  • injected storage failures.

Do not send deliberate high-volume traffic to systems you do not own or have permission to test. Set a budget and stopping conditions before the test begins.

7. Passing means more than “it did not crash”

A reasonable performance target might be 95% of ordinary requests under one second, 99% under three seconds, and unexpected 5xx errors below 0.1%.

Some outcomes should be zero-tolerance:

  • unauthorized data exposure;
  • unauthorized actions;
  • duplicate scoring;
  • lost committed data;
  • duplicate billing.

One occurrence is enough to stop and investigate.

8. A 9/10 test plan can still have a 0/10 execution score

A 324-item checklist and a prioritized set of 23 risks can be excellent test design.

Before running them, however:

  • test design: roughly 8.5–9/10;
  • executed evidence: 0/10.

That is not failure. It means the project has moved from “we do not know what to test” to “we know exactly where to try to break it.”

Now break your own test environment.

9. Then something else hits the limit: the AI subscription

Fast AI-assisted development moves another bottleneck into view.

Claude Max 20x costs $200 per month on the web as of October 5, 2026. The “20x” describes per-session capacity relative to Pro. Session limits reset every five hours, but Max also has a weekly limit shared across models.[5]

So 20x does not mean unlimited.

A full day of long, repository-heavy coding, repeated reviews, and agentic work can consume a large share of a weekly allowance. Exact usage depends on model, context, and workload, so the reliable place to inspect it is Settings > Usage.[5]

10. “Fable is expensive” survives contact with the numbers

On Max, Fable 5 and 5.1 are included, but Fable usage can consume up to 50% of the weekly limit and Anthropic says it draws down the allowance faster than other Claude models.[6]

Usage-based prices make the difference visible:

Model Input / 1M tokens Output / 1M tokens
Sonnet 5.5 $2 $10
Opus 5.5 $4 $20
Fable 5.1 $10 $50

Fable 5.1 therefore has 2.5 times the input and output token price of Opus 5.5. Anthropic reduced Fable 5.1 cache-read pricing and says typical Fable workloads are around 25% cheaper than Fable 5, with highly agentic workloads saving up to roughly 45%, but it is still the expensive option in absolute token pricing.[6][7][8]

Subscription quota accounting is not identical to API billing, so do not convert these prices directly into Max weekly minutes. The broader lesson is simply that the strongest model is not free just because it is inside a subscription.

11. Do not send every screw to a crane operator

A practical split is:

Sonnet 5.5: routine fixes, well-understood bugs, repetitive edits, UI cleanup, clearly scoped implementation.

Opus 5.5: unclear root causes, architecture changes, cross-file review, specification conflicts, high-stakes pre-release judgment.

Fable 5.1: the hardest work where the extra capability is worth the higher usage cost or faster quota burn.

That is not a ranking of intelligence. It is cost-per-task engineering.

A crane is impressive. A screwdriver job still does not need one.

12. AI does not remove bottlenecks. It moves them.

The old flow was:

idea → weeks of coding → testing.

The AI-assisted flow can become:

idea → rapid implementation → testing, operations, server quotas, and AI quotas all arrive at once.

So “I built the app in a day” can be true.

A more precise version is:

“In a day, I reached the stage where I can start trying to break it.”

That is a much better definition of progress.

References (9)

  1. Cloudflare DDoS Protection / Proactive defense: and https://developers.cloudflare.com/ddos-protection/best-practices/proactive-defense/ developers.cloudflare.com
  2. Cloudflare Workers Limits / Pricing: and https://developers.cloudflare.com/workers/platform/pricing/ developers.cloudflare.com
  3. Cloudflare D1 Limits / Pricing: and https://developers.cloudflare.com/d1/platform/pricing/ developers.cloudflare.com
  4. OWASP Web Security Testing Guide wstg.owasp.org
  5. Anthropic Max plan support.claude.com
  6. Anthropic Fable models on your plan support.claude.com
  7. Anthropic Claude Opus 5.5 anthropic.com
  8. Anthropic Claude Sonnet 5.5 anthropic.com
  9. Cloudflare Workers Rate Limiting API developers.cloudflare.com
Advertisement

Read this today

Each one answers a question readers of this article tend to ask next.

Browse all articlesMore on AI

Find other articles

All articles

Mendoi-chan

Who runs this site

Mendoi-chan

She turns friction at work and in everyday life into clear structure and practical next steps.