How we separate skill from luck
Any rule you can write down will produce a number when tested against history. The number by itself says almost nothing. This page describes how Arthora decides what a number is worth.
The problem with a single result
Take a rule that trades forty times over five years and finishes up twelve percent. That reads like evidence. It usually is not.
Forty trades is a small sample. Enter at forty random moments, hold for the same lengths of time, and you will get a spread of outcomes — some well above twelve percent, some well below. If the rule’s result sits comfortably inside that spread, the rule has told you nothing that rolling dice would not have.
Nothing about the rule’s reasoning changes this. A sensible-sounding rule and an arbitrary one face the same arithmetic.
The comparison we run
Every result is compared against a cohort of 500 variants generated from the result itself. Each variant enters at random moments but is matched on the things that are not in question:
- the same number of trades,
- the same distribution of holding periods,
- the same instrument,
- and the same costs.
Holding those constant leaves exactly one difference between the rule and the cohort: when it chose to act. Whatever the rule is claiming to know, that is where it has to show up.
The result is reported as a percentile. Sitting at the 62nd percentile means 38 percent of random-entry variants did better. That is a fact about where a number falls in a distribution, and we report it as one.
Why costs are applied to both sides
A comparison that charges the rule and not the cohort is not a comparison. Statutory charges, brokerage and slippage scale with how often you trade, so an uncosted cohort flatters exactly the rules that trade most — the ones where the question matters.
Both sides pay the same schedule: securities transaction tax, exchange charges, brokerage and a slippage model. Every figure Arthora shows anywhere is net of these.
Walk-forward, not fitted
A rule tuned until it fits the past will fit the past. To be worth anything it has to be decided before the period it is judged on, so results are measured walk-forward: parameters are fixed using earlier data and scored on later data the fitting never saw.
We hold ourselves to the same standard publicly. Arthora’s forecast ranges are scored against what actually happened, and the record is published rather than summarised.
When we refuse to produce a result
Sometimes the honest output is not a number. Arthora refuses, names the reason, and offers the nearest thing it can do instead, when:
- an input has less recorded history than the test window — the missing stretch would have to be invented, and an invented backtest is not a worse backtest, it is a different and fictional experiment;
- an input has stopped updating — measured against the clock, never against the most recent value we happen to hold;
- a rule is written around a specific price level and is asked to run on a different instrument, where it would produce a clean and meaningless curve rather than an error.
What a small sample can and cannot support
Below roughly thirty trades, sampling error dominates anything a rule could be detecting. A result over six trades is real in the sense that it happened; it supports no statement about what happens next. Arthora says so plainly on every result, and does not soften it as the number gets more flattering.