Skip to content

How a trading system is tested

Risk, plan and practice

How a trading system is tested

A rule that reads perfectly clearly to the person who wrote it can still be read two ways by anybody else, and a rule that can be read two ways cannot be tested, because two applications of it produce two different records. Testing a set of trading rules turns out to be mostly an exercise in writing them down precisely enough that the test is repeatable at all.

9 min read, Reviewed

What you will be able to do

  • Express a set of trading rules precisely enough that two people would apply them identically
  • Describe the sequence from historical review to forward testing on demo
  • Explain what sample size means for any conclusion drawn
  • Explain why costs must be included in any test

One sentence, two readers, two records 

Take a sentence of the kind that appears in most written plans: a position is opened on a pullback to the moving average, with a tight stop. Hand it, and the same chart, to two people who both understand it perfectly. One measures the pullback against a shorter average, the other against a longer one. One treats the condition as met the moment price touches, the other only once a period has closed beyond. One reads tight as a distance from the opening level, the other from the extreme of the pullback. Neither has misread anything. The sentence did not decide those questions, so each reader decided them differently, and the two records describe two different rule sets that happen to share a sentence.

That is the whole difficulty of testing, and it arrives before any data does. A test is a procedure that produces the same record when it is run again, and applying an ambiguous rule across a decade of history produces a decade of decisions resolved in the moment, which is a record of the person doing the resolving rather than of the rule.

Key term

Systematic trading
Systematic trading follows rules fixed in advance for entry, size and exit, so the same market data produces the same decisions whoever is watching the screen and however they feel about it.

Key term

Trading plan
A trading plan sets out in advance, in writing, which markets a trader deals in, how positions are sized, what defines an entry and an exit, and how results are reviewed.

The word system is used loosely enough to be worth pinning down. It does not mean automated and it does not mean complicated. A system here is a set of rules complete enough that the rules, rather than the reader, decide what happens at every moment they are consulted. One that meets that standard can fit on a page; one that does not can run to twenty and still not be testable.

What makes a rule testable 

A rule becomes testable when every question it raises can be answered yes or no at a stated moment, using only information that existed at that moment. That standard has a small number of parts, and a specification naming all of them can be applied by somebody who has never spoken to its author.

  • The universe and the timeframe. Which instruments the rules apply to, and which period length the conditions are evaluated on. A rule applied to whichever chart happens to be open is not the same rule twice.
  • The moment of evaluation. Whether a condition is checked at the close of a completed period or continuously within one. A condition true in the middle of a period and false at its close is two different observations.
  • The condition itself, in terms that can be computed. Every parameter named, and no word left that has to be interpreted. Strong, tight, near, clean and confirmed are all interpretation wearing the clothes of a specification.
  • The size, and how it is derived, whether from stop distance, from a volatility measure or from something else. A column of outcomes whose sizes were chosen by feel describes the sizing rather than the rules.
  • Both exits. What closes a position that has moved against, what closes one that has moved in favour, and the treatment of a position still open at the end of the span under review.
  • The exclusions. What the rules do around scheduled releases, at session boundaries, and when conditions appear at once on instruments that move together. An unstated exclusion is not the absence of a rule. It is a decision that will be made silently, later, under pressure.
Worked example. Illustrative figures, not YAL prices or terms.

One phrase, and the choices it leaves open

The phrase as written
a pullback to the moving average, with a tight stop
Which average, undecided
20 period, 50 period, 200 period
Which calculation, undecided
simple mean or exponentially weighted
Evaluated on which period length, undecided
5 minute, 1 hour, 1 day
Condition met when, undecided
price touches, or a period closes beyond
Tight measured from where, undecided
the opening level, or the extreme of the pullback
Tight meaning what distance, undecided
a fixed number of pips, or a multiple of a volatility measure
Combinations the phrase admits
3 × 2 × 3 × 2 × 2 × 2 = 144

The alternatives listed are assumptions of this block, chosen only to count the openings a familiar phrase leaves in itself. No combination of them is put forward here, none is described as preferable to another, and nothing whatever is stated or implied about what any of the 144 would produce. The count is arithmetic about a sentence, not about a market.

Restating a phrase this way is uncomfortable, for an informative reason. Much of what makes a written rule feel workable is the unwritten judgement surrounding it, and specification is the act of removing that judgement. Occasionally no fixed set of choices reproduces what the author believed they had been doing, and discovering that is a result of the exercise rather than a failure of it.

The sequence a test runs in 

Testing runs in an order, and the order matters because each stage costs less than the one after it and disqualifies work the next stage would otherwise waste.

  1. The rules are written and frozen, meaning recorded in a form that cannot be adjusted quietly once outcomes start appearing. A rule amended halfway through a test converts the test into a description of the amendments.
  2. The data is assembled and its provenance stated: which instrument, which span, which feed, and what one period on that feed represents.
  3. The rules are applied to that history, one period at a time, with everything after the current period concealed.
  4. Every outcome is recorded, including the occasions the rules declined to act and the occasions ambiguous enough to have required a judgement. An ambiguity found here sends the rules back to the first step rather than being resolved in passing.
  5. The same frozen rules are applied forward, in real time, on a demo account, where periods arrive one at a time and no part of the record exists yet.
  6. The two records are reviewed side by side, and the first question asked of them is whether the rules were applied as written.

That last question is what separates a test of a rule set from an assessment of its outcomes. Whether the rules were followed is a fact about the record and can be checked entry by entry. What a set of outcomes implies about outcomes still to come is a statistical question, taken up below, and it does not have the answer people generally want.

Applying rules to history without seeing the future 

Reviewing history is the cheap stage, and its one procedural requirement is the one most easily broken. Each decision has to be made with the rest of the chart concealed, because a chart displayed in full shows where price went next, and knowing that changes which conditions look like conditions. This is not a remark about anybody's discipline. A person reading a completed chart cannot unsee its right hand side, which is why the concealment is mechanical: the chart is advanced one period at a time, and each decision is written down before the next period appears.

The data itself carries assumptions worth stating rather than inheriting. Periods on a chart are assembled from a stream of quotes by a feed, and feeds differ in where they place an open and a close, in whether they record the bid, the ask or a mid, and in which quotes they include at all. A rule evaluated at a close is evaluated on a number another feed would have printed differently, so any record is a record of the rules as applied to one feed's version of the history. How quotes become periods is set out in the guide to quote feeds.

A second assumption concerns what is presumed to have happened at each recorded level, since a historical review conventionally assumes a position opened and closed exactly where the rules named. Real closes are not guaranteed at a named level: a market that gaps, or that moves faster than an instruction can be filled, closes a position wherever the next available price is, so a realised loss can be larger than the one the rules specified, and losses on a leveraged contract are not limited to the amount deposited. A record built on exact fills describes a friendlier instrument than the one that exists.

Trading CFDs and leveraged products involves a significant risk of loss and is not suitable for all investors. You could lose more than your initial investment. Ensure you fully understand the risks and seek independent advice if necessary.

Forward testing, where the future is genuinely absent 

Key term

Backtesting
Running a fixed set of trading rules over stored historical prices to record what that rule would have produced, which measures the rule against one past sample and nothing else.

A forward test applies the same frozen rules to periods that have not happened yet. Nothing about the rules changes. What changes is that the record is produced in the order and at the speed the market produces it, with the right hand side of the chart genuinely absent rather than covered over. A condition legible in review because the periods around it were visible is, forward, one period among others with nothing after it, recognised while it is still the last on the chart.

Key term

Demo account
An account running the same platform and the same quote stream as a funded one, in which every fill is produced by a simulator rather than obtained from a market.

A demo account is the conventional venue for that stage, and the line between what it reproduces and what it does not is worth drawing carefully. It reproduces the instrument list, the quotes, the order types, the arithmetic of an open position and the platform the rules would be operated on, which at YAL means MetaTrader 5. It does not reproduce a fill. An order on a demo account is matched against a simulation, so the level at which a position closes is a calculation rather than an event, and the gap between a simulated fill and a real one in a fast market appears nowhere in the record. It also does not reproduce the weight of a real loss, which is the input every behavioural failure runs through.

A record produced on a demo account establishes that the rules can be applied as written and carried out at the speed the market moves. It establishes nothing about what those rules would have returned against real fills, real financing and real money, and the size of that difference is not knowable in advance.

A test computed without costs describes nothing tradable 

Every position is opened and closed across a spread, so the level at which one opens is not the level at which it could be closed at that same instant. A commission, on the accounts that charge one, applies on the way in and again on the way out. A position held past the daily cut off carries a financing adjustment for every night it stays open. None of these are occasional. They apply to every position, in both directions, and they are the one component of a record that is certain in advance.

Their effect is not proportional to the size of anything else, which is what makes omitting them so distorting. Cost is charged per position and per night, so it scales with how often the rules act and how long they hold, not with how far price moved. Rules that act rarely and hold for weeks carry a small charge at the two ends and a large accumulated financing component; rules that act many times a day carry the opposite.

Worked example. Illustrative figures, not YAL prices or terms.

An assumed cost per position, applied across a column in both directions

Assumed cost per position, entry and exit combined
10.00
Positions in the assumed column
200
Total cost across the column
200 × 10.00 = 2,000.00
Adverse case, gross column
−1,500.00
Adverse case, net of the cost
−1,500.00 − 2,000.00 = −3,500.00
Favourable case, gross column
+1,500.00
Favourable case, net of the cost
+1,500.00 − 2,000.00 = −500.00
The same cost across 20 positions instead
20 × 10.00 = 200.00

Every figure here is an assumption of this block, chosen to keep the subtraction legible. None of it is a YAL cost, none of it is attached to a strategy, an instrument or an account, and the block says nothing about how often either arrangement occurs or about what any method produces. The adverse case is stated first and both are stated at equal prominence. Financing on positions held overnight is excluded from the assumed cost and would be added to it.

The subtraction is large enough to change the sign of a column, and it grows with the number of entries rather than with their size, so omitting it does most damage exactly where entries are most numerous. The charges an account carries are published rather than estimated, and a test built on the published figures is testing the instrument that exists.

What a number of recorded outcomes supports 

Key term

Expectancy
Expectancy is the average result per trade a set of rules produced over a sample of closed trades, combining how often it won with how much it won and lost.

Every conclusion drawn from a test is a statement about a sample, and the sample is the number of completed outcomes the rules produced, not the span of history reviewed. A decade of data for rules that act twice a month is a few hundred outcomes. A decade for rules that act twice a year is a list short enough to read in a minute, however impressive the calendar span sounds. The span is what it cost to collect the sample. It is not the sample.

Two further reductions are usually left out of the count. Recorded outcomes are not independent: entries produced within the same few weeks came from the same conditions, often on instruments that move together, so a run clustered in one stretch of the record carries less information than the same number spread across conditions that differed. And any parameter chosen after looking at the data has already spent part of the sample, because a threshold selected for producing the tidier record was fitted to that record, leaving only the portion never consulted to judge it by.

Worked example. Illustrative figures, not YAL prices or terms.

What a stated number of outcomes costs in calendar time

Assumed outcomes per month, infrequent rules
4
Months to accumulate 100 outcomes at that rate
100 ÷ 4 = 25
Assumed outcomes per month, frequent rules
20
Months to accumulate 100 outcomes at that rate
100 ÷ 20 = 5
Of those 100, an assumed number falling in one 3 month stretch
60
Distinct conditions the 100 were drawn from, on that assumption
closer to 3 than to 100

The rates and the clustering are assumptions of this block, attached to no method, no instrument and no account. The block states what a count costs in time and how clustering reduces it. It states nothing about what any number of outcomes establishes, and no number appearing here is described as sufficient.

That arithmetic is the part of testing that gets skipped, because a sample large enough to support much of anything takes an amount of time few people are willing to spend. The usual accelerations all buy quantity by spending independence: running the rules across many instruments at once produces entries that move together, shortening the timeframe so entries arrive faster draws them from a narrower slice of conditions, and passing over one stretch of history repeatedly produces a parameter chosen by the very data it is then judged on.

Why no rule set and no result appears here 

This lesson describes a procedure for testing and supplies nothing to test. That is a rule rather than an omission. A rule set published by a broker to the people trading through it would function as a recommendation whatever label sat above it, and this site makes none, for any technique. No result of any test appears anywhere on this site either: a figure describing how a set of rules performed would require a verified source, none exists, and a figure attached to a method is a claim about that method's effectiveness. What transfers between one reader and another is the procedure.

Where practitioners disagree 

The first argument is whether reviewing history is worth doing at all. One tradition holds that it is the only way to see many different conditions without waiting years for them, and that rules which cannot be applied consistently to history have been disqualified before any money is involved. Another holds that assumed fills, a single feed and the impossibility of unseeing a completed chart make a historical record so unlike a live one that its main product is confidence rather than information. Both camps agree on one point: a stretch of history can be used up, because each pass over it with different parameters spends some of the independence that made the next pass meaningful.

The second argument is how long a forward phase runs. One convention sets a fixed number of recorded outcomes, on the reasoning that the sample is what matters and calendar time is only its price. Another sets a fixed span of time, since a count reached quickly was reached inside one set of conditions while a span at least allows the conditions to change. A third holds out for several distinguishable conditions, and is objected to on the ground that nobody agrees where one condition ends and the next begins.

The third argument is whether a demo phase says anything about the person operating the rules. One camp holds that the mechanics are identical, that the mechanics are what is under test, and that somebody who cannot follow a rule set on a demo account has learned it cheaply. Another holds that nothing involving simulated money resembles losing real money, so a demo phase tests the rules and leaves the operator untested. Both agree on the limit: a demo record cannot establish what fills would have been, and the behavioural failures this module closes on are the ones no test of this kind reaches.

In summary 

  • A rule is testable when two people reading it would act identically. That requires the universe, the timeframe, the moment of evaluation, the computable condition, the sizing, both exits and the exclusions all to be stated, with no word left that has to be interpreted.
  • The sequence is: freeze the rules, state the data and the feed, apply them to history one period at a time with the future concealed, record every outcome including the ambiguous ones, then apply the same frozen rules forward in real time on a demo account. A demo reproduces the mechanics, not the fills and not the weight of a real loss.
  • The sample is the number of completed outcomes, not the span of history reviewed, and it shrinks further when outcomes cluster in one stretch of conditions or when parameters were chosen after seeing the data. Every shortcut to a larger count spends independence to get it.
  • Spread, commission and overnight financing apply to every position in both directions, and they scale with how often the rules act rather than with how far price moves, so a record computed without them is not a record of anything tradable.

Get started

Open your account in four steps.

A clear path from sign-up to your first trade, in four steps.

No depositNo documents

  1. 01/ 04step 1 of 4

    Register

    A few details to get started.

    No deposit to open

  2. 02/ 04step 2 of 4

    Verify

    Confirm your identity, securely.

    ID and proof of address

  3. 03/ 04step 3 of 4

    Fund

    Add money by bank transfer or card.

    From $0

  4. 04/ 04step 4 of 4

    Trade

    Go live on the platform you already know.

    MetaTrader 5

Cookies on this site

Some cookies are needed to make the site work. With your permission we also use analytics cookies to see which pages are read, so we can improve them. You can change your choice at any time.