Risk note. Historical statistics describe past behaviour under one set of assumptions. No metric on this page indicates future results, and leveraged products carry substantial risk to your capital.
The Direct Answer
R-squared tells you how tightly an equity line tracks a fitted linear trend across the observations tested. Higher readings mean the points cluster tightly around that line. Lower readings mean the path wandered further from it.
That’s all it tells you. Specifically, it does not tell you whether the trend rises or falls, whether declines along the way were survivable, whether the underlying logic is sound, or whether anything similar will happen next.
Which matters because the 2019 version of this article treated a higher reading as straightforwardly better, and that framing was wrong in a way worth correcting properly rather than quietly.
Two Definitions I Got Wrong
Both errors came from the same source: describing the statistic visually rather than mathematically.
“Zero is when the equity line doesn’t touch the fitted line.” Not accurate. The calculation has nothing to do with whether points physically sit on the line. It compares the size of the residual differences between observed values and fitted ones against the total variation in the data. A reading near zero means the linear fit explains almost none of that variation, which is a statement about residual magnitude, not about contact.
“A perfect 100 requires only two observations.” Also wrong. Any number of points can produce a perfect fit, provided every one of them sits exactly on the fitted line. Two points always produce a perfect fit because two points define a line, but that’s a special case rather than the rule. A perfectly linear equity line across four hundred closed positions is extremely unusual in practice, though nothing about the mathematics forbids it.
Apologies for propagating both. The second in particular has been repeated back to me by readers, which is how bad explanations spread.

What the Calculation Is Doing
Take your equity line. Fit a trend through it using linear regression. Then measure how much of the variation in the observed values that fitted trend accounts for.

The result is expressed either from 0 to 1 or from 0 to 100, depending on the platform, and it answers exactly one question: how well does a linear fit describe this data?
| Reading | What it means |
| Near 100 | Observations sit close to the fitted trend |
| Middle range | Substantial deviation, though a linear pattern remains visible |
| Near zero | The linear fit explains very little of the variation |
Note what’s absent from that table. Direction. Magnitude. Risk. A statistic measuring fit quality cannot report on any of them, because those aren’t the question it asks.
The Losing Strategy Problem
Here’s the example that should have been in the original article.
Imagine an equity line declining steadily, losing roughly the same amount every month, with almost no deviation from that downward path. Fit a trend through it and the fit is superb. R-squared near 95.
Consistently losing money is still consistently losing money. The statistic reported linearity accurately, and linearity was never the thing you cared about.
| Example | Net result | R-squared | Reading |
| A | +20% | 92 | Positive outcome along a fairly linear path |
| B | −20% | 95 | Very linear, and consistently unprofitable |
| C | +25% | 48 | Made money, though the route was irregular |
Example B is why this metric must never be read alone. Example C is the subtler lesson: a low reading does not condemn a strategy. Irregular paths can still produce acceptable outcomes, and rejecting C purely on smoothness would discard the best net result in the table.
Pair the reading with direction, always. The slope of that fitted trend tells you which way things went; the fit quality only tells you how tidily.
What It Cannot Establish
- Whether the strategy made money: Direction lives in a separate number.
- Whether declines were tolerable: A path can track a trend reasonably while containing a decline you could never have sat through.
- Whether stagnation periods were short: Related to the above, and separately measured.
- Whether the logic is sound: Fit quality says nothing about why entries happen.
- Whether results will persist: This describes history only.
- Whether the sample is adequate: Twelve observations can produce a beautiful fit that means nothing.
That last point deserves its own section.
Sample Size Changes Everything
A high reading calculated from few observations is close to meaningless, since fitting a line through a handful of points is easy and tells you almost nothing about the process generating them.

Minimum thresholds I use in practice: at least 10 observations before the number is worth displaying at all, and 300 or more before I take it seriously. One example from my own collection showed a count of 333 alongside its reading, which is roughly the territory where the statistic starts carrying information.
Those figures are working preferences rather than validated standards. I have not tested whether 300 outperforms 200 or 500 as a cutoff, and neither has anybody else as far as I know.
Balance Line or Equity Line?
The original article used “equity line”, “balance chart”, and “balance line” interchangeably. They aren’t the same thing.
| Series | What it records |
| Balance | Changes only when positions close and results are realized |
| Equity | Includes the floating value of anything currently open |
Calculating from realized balance produces a different reading than calculating from mark-to-market equity, sometimes substantially, because a strategy holding losing positions for weeks looks far smoother on a balance series than on an equity one.
Any platform reporting this metric should state which series it uses. Ask, if the documentation doesn’t say.
What Sits on the Horizontal Axis
Rarely disclosed, and it changes the interpretation completely.
The calculation needs something on that axis, and platforms choose differently:
- Position number: Measures consistency from one closed position to the next.
- Bar number: Ties the reading to chart periods rather than activity.
- Calendar time: Includes stretches where nothing traded at all.
- Daily or weekly sampling: Smooths short-term variation before fitting.
A strategy trading twice yearly looks very different measured by position number than by calendar time, since the calendar version counts every quiet month as data while the position version ignores them.
Settings worth disclosing whenever you publish a reading:
| Setting | What to state |
| Series fitted | Balance or equity |
| Independent variable | Position number, bar number, or time |
| Sampling frequency | Every position, daily, weekly |
| Scale | Account currency, percentage, or normalized |
| Display range | 0 to 1, or 0 to 100 |
| Deposits and withdrawals | Included, excluded, or adjusted for |
| Open positions | Counted or ignored |
Two people reporting different readings for the same strategy usually differ on one of those rows rather than disagreeing about mathematics.
Deposits and Sizing Distort It
An equity line can appear tidier or steeper for reasons unrelated to strategy quality:
- Capital added or withdrawn mid-period.
- Position sizes increasing as the account grows.
- Compounding, which curves an otherwise linear path upward.
- Leverage changed partway through.
- Risk per position adjusted after a good stretch.
- Several strategies running in one account simultaneously.
Compounding deserves particular attention. A strategy producing steady percentage growth generates an exponential account path, which fits a linear trend poorly and scores badly despite being exactly what you’d want. Fitting the logarithm, or working from percentage returns rather than currency amounts, handles this properly.
Reading It Alongside Everything Else
| Metric | What it measures | Where it falls short |
| R-squared | Linearity around a fitted trend | Silent on direction and risk |
| Net result | Total historical gain or loss | Ignores the path taken |
| Maximum drawdown | Largest peak-to-trough decline | Says nothing about consistency |
| Profit factor | Gross gains against gross losses | Unstable on small samples |
| Position count | Number of observations | Quantity is not quality |
| Maximum stagnation | Longest period without a new high | Doesn’t capture loss depth |
| Out-of-sample result | Behaviour on unseen data | Still depends on the chosen sample |
My earlier claim that a higher reading means fewer large declines and less stagnation was too strong. The two often move together, and a tidier path usually does contain shallower dips, but “usually” isn’t “necessarily”. A line can track its trend acceptably while containing one decline that would have ended your account. Measure drawdown separately. Always.
Interpreting Combinations
| Pattern | Reasonable reading | What it doesn’t establish |
| High reading, rising trend | History followed a fairly consistent upward path | Future results or acceptable risk |
| High reading, falling trend | Losses accumulated consistently | Anything usable |
| Low reading, positive return | Made money via an irregular route | That the approach is flawed |
| Low reading, negative return | Unprofitable and erratic | Why it failed |
| High in-sample, low out-of-sample | Historical tidiness didn’t survive unseen data | The cause of the deterioration |
| Similar across several periods | Linearity was reasonably stable in testing | That the edge persists |
Row five is the one worth watching for. Selecting strategies by this metric on development data, then finding the reading collapses on data you never touched, is a signature of fitting rather than finding.
Optimizing for Smoothness Has a Cost
Generation software lets you search for candidates by this metric, apply it as an acceptance threshold, or optimize existing candidates against it. All three work, and all three carry the same risk.
Screening thousands of generated candidates and keeping those with the tidiest historical paths is a selection process. Selection produces winners by chance when the pool is large enough, and a strategy that happened to produce a linear result across your test period may have nothing repeatable behind it.
Worse, smoothness is precisely the property that curve-fitted results tend to display, since a strategy fitted closely to historical data reproduces that data neatly by construction.
So use it as one filter among several rather than as the objective. Out-of-sample confirmation, Monte Carlo stress testing, walk-forward analysis, and forward testing on a demo account all remain necessary afterward.
How This Works in EA Studio
Disclosure first: I have a commercial relationship with Expert Advisor Studio and sell courses covering it. What follows describes one implementation rather than the only way to obtain this statistic.

The reading appears in the backtest output panel, and which metrics display there is configurable through the settings menu, so you can swap in position count or the win-loss ratio depending on what you’re assessing.
Three places it can be applied:
- Generator search method: Instead of ranking candidates by net balance, rank them by this reading. The generator then favours tidier historical paths.
- Acceptance criteria: Set a floor, 70 for instance, and candidates falling below it never reach your collection.
- Optimizer target: Generate by net balance, which is what most people want since results ultimately matter more than tidiness, then optimize the survivors against linearity afterward.
Readings from my own collection span a wide band: 77.35 on one candidate, 88.09, 89.16, and 95.99 on others produced with this metric as the optimization target. One showed 92.64 alongside maximum stagnation of 13.5 percent.

I described 92.64 as “very good” in the original article. Compared with what, though? I have never published a distribution of readings across generated candidates, nor tested whether higher-scoring strategies subsequently outperformed lower-scoring ones out of sample. Until somebody does that work, treat these numbers as workflow preferences rather than standards.
Common Mistakes
- Reading it without direction: The single most consequential error, and the one my earlier article encouraged.
- Trusting it on tiny samples: Ten observations prove nothing.
- Assuming it bounds drawdown: Separate metric, separate question.
- Optimizing for it exclusively: Selects for the appearance of consistency.
- Comparing across different calculation settings: Position-number and calendar-time readings aren’t comparable.
- Ignoring compounding: Penalizes exactly the growth pattern you want.
- Treating a threshold as validated: Mine aren’t, and I doubt anyone else’s are either.
A Practical Checklist
Before letting this statistic influence a decision:
- Confirm which series the platform used, balance or equity.
- Confirm what sits on the horizontal axis.
- Check the observation count; below a few hundred, discount heavily.
- Read the direction of the fitted trend before anything else.
- Read maximum drawdown and stagnation separately.
- Compare the in-sample reading against out-of-sample.
- Adjust for deposits, withdrawals, and compounding if present.
- Treat any threshold as your own working preference.
Frequently Asked Questions
Is this the same as the correlation coefficient?
Related but distinct. For simple linear fits, R-squared equals the square of that statistic, which discards the sign in the process. That underlying figure runs from negative one to positive one and tells you direction as well as strength; squaring it collapses everything into a range from zero to one where direction disappears entirely. That loss of sign is exactly why a declining equity line can score highly, and why direction must be read from a separate figure.
What about adjusted R-squared?
The adjusted version penalizes explanatory variables that don’t improve the fit meaningfully, which matters in multiple regression where researchers might otherwise add predictors indefinitely. For a single-predictor equity line fit, the adjustment makes little practical difference, so most trading platforms report the unadjusted figure. Worth knowing the distinction exists if you move into more complex modelling, since the plain version rises mechanically as variables accumulate.
Does a higher reading mean lower risk?
Not directly, though the two often appear together. Tidier historical paths usually contain shallower declines, so the association is real, but it’s a tendency rather than a guarantee. A single severe drawdown can sit inside an otherwise linear path without dragging the fit statistic down much, particularly across a long sample. Risk deserves its own measurement: maximum decline, longest stagnation, and worst losing sequence each answer questions this metric cannot.
Can I calculate it outside a trading platform?
Yes, easily. Export your closed position history, put the running balance in one column and the position number in another, then use a spreadsheet’s RSQ function or a few lines of Python. Doing it yourself has an advantage: you control the settings, so you know precisely which series and which axis produced the figure. That transparency is worth more than a number appearing in a panel whose calculation method is undocumented.
How does it compare to the Sharpe ratio?
Different questions entirely. That ratio relates average excess return to volatility, producing a risk-adjusted performance figure that incorporates both direction and dispersion. R-squared only describes how well a linear model fits, ignoring magnitude completely. A strategy could score highly on one and poorly on the other in either combination. Neither replaces the other, and reading both alongside drawdown gives a considerably fuller picture than either alone.
Should I reject strategies below a certain reading?
Only if you’ve decided what the threshold is doing for you. Screening out erratic historical paths is defensible when consistency matters for your own tolerance, since sitting through a jagged equity line is genuinely harder than a steady one. But rejecting a profitable irregular strategy in favour of a smoother marginal one optimizes for comfort rather than outcome. Decide which you’re prioritizing, and be honest about it.
Disclosure: The author has a commercial relationship with Expert Advisor Studio. Content here is educational and does not constitute a recommendation to trade or to purchase any software. Historical statistics describe past conditions only.

Petko Aleksandrov

