Where Does the Edge in AI Algorithmic Trading Actually Come From? Price Prediction, Order Books, Low-Leverage Tests, and the Power of Running More Experiments

Share this article
Advertisement
Advertisement

If you want to build an AI trading system, the obvious idea is simple: predict the next price.

That is also where the trouble begins.

Prices are noisy. The more configurations you test, the easier it becomes to discover a strategy that looks brilliant only in historical data. In scalping, even a mildly accurate forecast can be destroyed by fees, bid-ask spread, latency and slippage.[1][2]

The practical conclusion is this: the most useful AI in trading may be less of an oracle and more of an experiment engine—searching many candidates, classifying market regimes, measuring execution risk and killing fake edges quickly.

This idea became concrete after hearing about a nearby developer who had spent years working on AI-driven automated trading. Public third-party analytics showed many real, automated trading accounts.

The interesting question was not merely, “Is this person good?”

It became: Should we judge only the winning accounts? What remains after leverage is stripped away? Should AI predict prices, or should it read order flow? Can an individual compete in millisecond scalping?

Eventually the same logic even explained why an AI-powered article factory can produce huge volume and still suffer from a terrible defect rate.

Trading and content production both become a loop of experiment → inspect → reject defects → rerun.

1. The first rule: a real edge survives when leverage is removed

An edge is a repeatable advantage that retains positive expected value after costs and risk.

A spectacular return alone is not enough.

At minimum, examine:

  • maximum drawdown;
  • effective leverage and actual exposure;
  • returns after fees, spread and slippage;
  • performance on unseen periods and other markets;
  • behavior during stress events;
  • and how many failed strategies were tested before the winner appeared.

Leverage can amplify returns. It can also amplify a weak signal until it looks heroic.

What matters is the part that still looks interesting at low leverage.

Leverage is an amplifier, not vocal talent. If the song disappears when the amplifier is turned down, perhaps the amplifier was the main performer.

2. Read the entire research shelf, not only the winning account

Myfxbook offers verification for public trading accounts. “Track Record” indicates that the history shown matches the broker-side trading history, while “Trading Privileges” confirms control of the account.[3]

That is useful evidence, but it does not prove that an AI is sophisticated, that a strategy will remain profitable, or that its edge is causal.

In one public profile examined for this research, real automated systems from the same development line included one showing +163.46% gain with 20.90% maximum drawdown, while another showed -99.90% with 99.93% drawdown.[4]

Looking only at the +163% account creates a legend.

Looking only at -99.9% creates a disaster story.

Research requires looking at the distribution.

The more configurations a researcher tests, the greater the danger of finding a lucky backtest by chance. This is a central problem in the literature on backtest overfitting.[1]

Also, an account label such as “1:2000 leverage” does not prove that every trade used 2000x effective exposure. Account leverage settings, margin usage and actual gross exposure are different quantities.

So “that leverage is insane” is only the opening joke. The real question is: how much risk was actually deployed?

3. A winning public account does not prove the developer is rich

Trading-account performance cannot reliably reveal a developer's personal wealth or lifestyle.

A test account may contain research capital. Profits may have been withdrawn. Corporate and personal assets are separate.

Homes, cars, clothing and family appearance are especially poor wealth estimators.

A person can live in an old house and own substantial assets. A person building advanced AI is under no obligation to live in a glass tower.

Evaluate technology with technical evidence.

Do not backtest somebody else's wallet.

4. “I joined a ranking” requires checking what kind of ranking it was

Myfxbook has public “Most Popular Systems” lists and a separate “Contests” section.[5]

Therefore, a claim such as “I joined a ranking for visibility” could refer to a popularity list, a formal trading contest, or another platform entirely.

In this investigation, public systems were verifiable, but the specific contest implied by word of mouth could not be independently confirmed.

That distinction matters.

Fast research also creates fast misidentification. Do not upgrade “probably this” into “confirmed.”

5. Instead of predicting the next price, let AI classify the market state

For short-horizon trading, a more useful target can sometimes be the state of the market rather than the exact next price.

Possible questions include:

  • Is displayed bid depth stronger than ask depth?
  • Which side dominates executed trades?
  • Is the spread widening or narrowing?
  • Are visible orders persistent or rapidly cancelled?
  • Is volatility high or low?
  • Is the regime trending, ranging or dislocated?
  • Is spot, futures, another venue or a related asset moving first?
  • Does the expected move remain profitable after execution costs?

Research on Order Flow Imbalance—the asymmetry in incoming buy and sell flow—shows that order-flow structure can carry information about short-horizon price formation.[6]

Research also documents lead-lag relationships across related instruments using price, liquidity, book imbalance and volatility measures.[7]

But the order book is not a tablet of truth. Visible orders can disappear. Orders can be placed with little intention of execution. A robust system should examine additions, cancellations and actual trades—not only a static snapshot.

The useful AI question may therefore be less “Where will price be?” and more “What regime are we in, is this signal still valid, and will costs eat it?”

6. Scalping and millisecond scalping often die in execution, not prediction

As holding periods shrink to seconds or milliseconds, execution becomes part of the model.

You must care about:

  • API and network latency;
  • time from signal to exchange arrival;
  • queue position;
  • probability of limit-order fill;
  • spread paid by market orders;
  • maker/taker fees;
  • adverse price movement after a signal;
  • and market impact.

Research has shown that predictive information from top-of-book imbalance can dissipate quickly as latency rises.[2]

An individual trader therefore should not automatically choose a battlefield where the main contest is who reaches the exchange a microsecond earlier.

Move to a slightly slower horizon. Trade less. Subtract costs first. Combine lead-lag structure with regime classification.

“AI became smarter, so let us fight at 0.001 seconds” is a remarkably muscular use of intelligence.

7. The scariest AI trading failure is intelligent overfitting

Stronger AI can make backtesting more dangerous because it can test thousands or millions of variants faster than a human.[1]

A serious search process should therefore include:

  1. out-of-sample periods;
  2. multiple rolling time windows;
  3. fees worse than the observed baseline;
  4. slippage stress;
  5. execution delays;
  6. missed fills;
  7. crisis periods;
  8. parameter perturbation;
  9. a record of the total number of strategies tried;
  10. and small live deployment at the end.

AI is not only a strategy generator. It is also a machine capable of mass-producing convincing false positives.

The quality of AI research is determined as much by defect detection as by generation.

8. Five questions worth asking an experienced automated-trading developer

8-1. What exactly does your AI predict?

Price, regime, order flow, volatility, entries, exits or risk?

8-2. Which short-horizon data are still worth studying for an individual?

Order book, trades, volume, spot-futures relationships, cross-venue signals or cross-asset lead-lag?

8-3. Which edges survive low leverage?

Ask about structures that retain expected value when effective leverage is reduced.

8-4. Why do strategies win in backtests and die live?

Overfitting, fees, slippage, fills, liquidity changes and regime shifts are often more valuable topics than the secret winning rule.

8-5. If you started from zero today, what market, horizon, data and AI target would you choose?

These questions sound sophisticated not because they contain jargon, but because they come after actual experimentation and failure.

“Give me a profitable bot” is a beginner request.

“Direct price prediction was weak; what structure should I test next?” is a research conversation.

The question itself becomes a calling card.

9. Crypto exchange access depends on residency rules, not clever workarounds

Access for Japanese residents has changed substantially.

Bybit stopped accepting new registrations from Japanese residents from October 31, 2025.[8]

Bitget stopped new registrations from Japanese residents after August 3, 2026 and announced a Close-Only restriction from November 1, followed by forced closure of remaining positions from December 31 for affected residents.[9]

Binance operates a Japan-specific platform, Binance Japan.[10]

The right response is not to hunt for regulatory loopholes. Check residency, identity verification, platform terms and local law together.

If a person genuinely moves abroad, address and tax-residency records should be updated accurately. Faking residency through a VPN can create account and legal risks.

Regulation is part of system architecture.

10. Using a new frontier model on launch day: the real edge is low action latency

OpenAI announced GPT-6 Astra on September 3, 2026, describing improvements in coding, research, computer use and complex multi-step work. Its API documentation lists a 1.05-million-token context window.[11][12]

The interesting behavior is not admiring a new model. It is deploying it immediately into a real research loop.

Test price prediction.

If it fails, test the order book.

Add trades.

Test lead-lag.

Add fees.

Add latency.

Watch it die.

Build the next one.

The advantage is not supernatural first-shot accuracy. It is extremely low delay between hypothesis and experiment.

AI-era productivity can be viewed as:

quality of hypotheses × experiment speed × parallelism × defect-rejection speed.

11. There is a difference between “using AI” and building a one-person research lab with AI

Basic AI use is becoming normal: search, rewriting, summarization and coding.

The next level is a closed loop:

hypothesis → data → AI exploration → stress tests → live comparison → expert questions → next experiment.

At that point, AI is a research assistant rather than a convenience feature.

Run research, coding, analysis and publishing in parallel and one person begins to resemble a tiny research institute plus editorial desk.

Outsource only the in-person interview and you get the absurd newsroom model:

Interview: friend / Analysis: AI / Article: author / Author: staying home.

The self-proclaimed reporter title becomes questionable. The workflow design remains excellent.

12. Do not type people by MBTI; compare their iteration behavior

Long-term obsession with AI, automation and system design may tempt observers to assign a personality type.

That is weak evidence. MBTI-style categories are self-report classifications and should not be imposed on another person from the outside.

Behavior is more informative:

  • staying with one technical theme for years;
  • building many versions;
  • breaking them in live conditions;
  • continuing after failures;
  • testing new AI models immediately;
  • caring about failure modes, not only wins.

If two people seem similar, perhaps their iteration loops are similar—not necessarily their four-letter labels.

13. An AI article factory has the same problem: throughput is no longer the bottleneck, yield is

Once AI can generate large volumes of articles, the next problem is not “write more.”

Suppose 100 drafts enter a pipeline:

  • generated: 100;
  • all languages complete: 80;
  • quality passed: 60;
  • formally accepted: 45;
  • published: 40.

Final yield is 40%.

Doubling generation to 200 may simply double the defect pile.

The better metric is First Pass Yield: the percentage that reaches the final stage without rework.

Typical article-factory defects resemble trading-system failures:

  • missing languages;
  • worker rules drifting apart;
  • generated content never being formally accepted;
  • aggregates not updating;
  • quality gates rejecting output;
  • handoff stages failing;
  • jobs stopping because of conflicts or freeze logic.

The next star employee is therefore not “the writing AI.” It is the AI defect-analysis supervisor.

A stronger model can help trace dependencies across stages, but a smarter model alone cannot fix a badly designed factory. Canonical sources, quality gates, retries, conflict prevention, aggregation and deployment checks still have to be engineered.

Trading and content production converge on the same lesson:

AI is not a magic profit machine. It is equipment for faster experimentation and quality control.

14. Conclusion: the biggest AI-era edge may be how many high-quality failures you can run

The goal in AI trading is not a beautiful backtest, maximum leverage or an oracle that speaks with confidence.

The goal is the stubborn component that survives:

costs, latency, unseen periods, low leverage and live trading.

Order books, trades, order flow, lead-lag and regime classification are all worth exploring. None is an edge by itself. Execution costs and competitors must be included.

AI's deepest advantage may be increasing the number of serious experiments you can run.

Build.

Break.

Diagnose.

Change conditions.

Run again.

There may be no one-shot automated trading system and no one-shot perfect content factory.

AI-powered productivity is not about being correct every time.

It may be about running 100 legitimate failures while somebody else is still debating experiment number 10—and already starting number 101.

FAQ

Is direct price prediction with AI useless?

No. But exact-price targets are often noisy and vulnerable to overfitting. Regime, order flow, volatility and net expected value after costs can be easier targets to validate.

Can an order book alone produce a scalping edge?

It can contain short-horizon information, but cancellations, spoofing, queue position, latency and costs may erase it. Test order events and actual trades, not only static depth.[2][6]

Does “Real” and “Automated” on Myfxbook mean a system is safe?

No. It is useful evidence about account type and automation, not a guarantee of future returns, safety or AI quality.[3]

Are high-leverage account results meaningless?

No. But distinguish the account leverage setting from effective exposure, and normalize results for drawdown and actual risk.

Will a stronger AI model automatically finish an automated-trading system or article factory?

No. Better models improve exploration and diagnosis. Weak data, evaluation, costs and operational controls simply allow a smarter system to manufacture defects faster.


Sources

  1. Bailey, Borwein, López de Prado & Zhu, “The Probability of Backtest Overfitting papers.ssrn.com
  2. Stoikov & Waeber, “Reducing Transaction Costs with Low-Latency Trading Algorithms papers.ssrn.com
  3. Myfxbook Help, “Verification myfxbook.com
  4. Myfxbook public profile, TestSamurai30Ai, systems list myfxbook.com
  5. Myfxbook, “Most Popular Systems” and Contests navigation myfxbook.com
  6. Anantha & Jain, “Forecasting High Frequency Order Flow Imbalance arxiv.org
  7. Schmidt, Cestonaro & Bender, “Lead-Lag Relationships in Market Microstructure papers.ssrn.com
  8. Bybit, “Bybit to Pause New User Onboarding in Japan bybit.com
  9. Bitget, “Important Notice for Japanese Residents” / FAQ — ; https://www.bitget.com/support/articles/12560603890271 bitget.com
  10. Binance, “Introducing Binance Japan: A Platform for Japanese Residents binance.com
  11. OpenAI, “GPT-6 Astra: A new generation of intelligence openai.com
  12. OpenAI API, “GPT-6 Astra Model developers.openai.com
Advertisement

Find other articles

All articles

Mendoi-chan

Written by

Mendoi-chan

She turns friction at work and in everyday life into clear structure and practical next steps.

About
Advertisement

Latest articles

  1. 1Was an 18-Hour Sleep Day Recovery Sleep? How to Understand Long Sleep and Changes in Dreams
  2. 2Do You Really Need to Apologize for Not Giving Your Parents Grandchildren? Sometimes an Adult Child Coming Home for Dinner Is Already a Big Deal
  3. 3The Day a 40-Year-Old VTuber Became a “Digital Community Center”: Age Does Not Always Kill Demand—Sometimes It Changes Its Shape
  4. 4I handed senior-level engineering to an AI agent from my phone—and the move finished first
  5. 5How AI Article Automation Turned Into an Autonomous Factory in About a Week: One Ultra Punch, Level 6, and Why Level 7 Can Wait

You may also like

Advertisement