The Waving Flag: ADLG: Super Armies, The Final Solution

Friday, 25 September 2026

ADLG: Super Armies, The Final Solution

Summary

Since 2018 I have always answered no to the question are there any "super" armies in Art de la Guerre (ADLG)? After the thorough statistical testing described here that's still the case, but for entirely different reasons.

Until today I thought, given a high enough number of games, any variation in army performance (efficacy) could be ascribed to the inherent strength of the list and nothing else. That's not the case.

Chance genuinely evens out, and what is left behind is real and not chance - but that leftover is a tangle of list strength, player skill and popularity. The statistical tests only measure its size, but cannot pull the tangle apart. This applies to all the data I have available. More data will not help.

The answer is definitive, but it's case unproven. Therefore this post is the final piece to look at the issue of "super" armies in ADLG. Time to move on.

Those of a statistically nervous disposition read on at their peril.

Introduction

This is the fifth time I've written about this. I was (am) interested in whether there are armies that offer all players an advantage. This is my definition of a "super"" army. Not whether an army does well in a few high profile competitions, but whether there's something inherent in the list that helps everyone most of the time.

My last piece on the subject was only published days ago, but I'd identified something about armies with lots of shooting power that kept nagging at me. Enough to seek better statistical tools.

I have always struggled to analyse this data to my satisfaction not least because my statistical knowledge is wholly inadequate for the task. Now, thanks to claude.ai, I have finally been able to analyse the data properly.

Background

The data I use comes from the ADLG statistics page. I separate the data in to sets for V3 & V4 armies. Both data set are large containing over 40,000 "player games" each (every game results in two player games - one for each player).

Plotting army efficacy against games played has always shown a clear trend: efficacy tends towards 50% as the number of games increases:

Efficacy v Player Games for ADLG V4

The online dashboard shows the same pattern for V3 & V4.

My underlying hypothesis (more of this later) was that as the number of games increases things like; list choices, player skill, and dice luck should even out leaving just the inherent differences (if any) between the lists. From the chart this looks to be the case but looks can be deceiving.

Questions raised

Firstly, in the chart the most popular army (127, Nikephorian Byzantine) shows a persistent value of 57.28% with a large number of games. Therefore it is a candidate for a super army. However, comparison with the pooled standard deviation (all 300 armies) suggested it wasn't a significant difference (I was told later this was an inappropriate statistical test).

Secondly, I spotted a possible trend in the armies with more than 500 games: those with a lot of shooting units seem to do slightly better (Nikephorians, Samurai, French Ordonnance, Ottomans). When I looked at this subset in isolation some differences were indeed statistically significant.

So, depending on how I slice the data I get different results. Not good. Plus there was nothing I could do would prove my underlying hypothesis. Most unsatisfactory. I must have been doing something wrong.

Enter claude.ai.

Full statistical analysis

The aim was simple: test my underlying hypothesis and see if army 127 was significantly different. To spare you the gory details here are the results from the AI:

Hypothesis tested: as games increase, luck, player skill and list choice even out, so efficacy tends to 50%, leaving only inherent list differences (if any).

  1. Luck evens out, but efficacy does not converge to 50%. More games do shrink the random swing in a result, exactly as the hypothesis predicts. But what a result converges to is each army's own true strength, not 50% — for most armies that is close to 50%, but not for all.
  2. Real, persistent deviations from 50% exist, and they are not chance. The spread left over once luck is subtracted out (the persistent spread) is about 4.5 points in V4 and 5.0 points in V3. The heterogeneity test rules out luck as the explanation, with p-values as small as 10⁻³⁹ to 10⁻⁵³.
  3. A worked example (army 127): raw efficacy 57.3%, with a 95% confidence interval of 54.5% to 60.1%. Its shrunk estimate — the best single guess at its true strength after allowing for some of that gap being luck — is about 56.6%.
  4. Chance is effectively ruled out for such armies, but the cause behind the deviation is not identified. A z-score of 5.1 (p ≈ 4×10⁻⁷) makes luck alone an implausible explanation for army 127's result. But the test cannot say whether that gap comes from list design, player skill, popularity, or some mix of all three.
  5. More games sharpen the estimate, not the explanation. Extra games would keep narrowing an army's confidence interval, with steeply diminishing returns once it already has hundreds of games. But win/draw/loss totals never record why a game was won, so no volume of extra games can, on its own, separate list strength from player skill.
  6. The hypothesis is proven half right. "Luck evens out" is confirmed. "A real difference remains, if any" is also confirmed — the "if any" resolves to yes. But the claim that what remains is purely inherent list differences is not established: player skill and list popularity are still tangled up inside that same persistent number. One finding points this way directly — popular armies (more games played) score slightly better on average, hinting that some of what looks like list strength may be player selection instead.

In one line: luck genuinely evens out, and what is left behind is real and not chance — but that leftover is a tangle of list strength, player skill and popularity, and the tests can only measure its size, not pull the tangle apart.

Closing remarks

There you have it: my hypothesis was right in part but I really can't tell if there any "super" armies. It's case unproven one way or another.

There are significant differences between armies but it's impossible to say if the cause(s) are inherent in the lists or not. Player skill and list choices cannot be disentangled with the data available.

What makes it worse is that I never stood a chance with the data I had. So, after eight years(!), here endeth the search.

Supporting documents

There's lot more to the analysis than the above including lists of armies that do well and badly under both V3 & V4. See:

The second document is well worth reading as there're lots of good points for the wargamer, but it's too long for a blog post. I may return to it one day soon.

Related posts

3 comments :

Jonathan Freitag said...

Interesting results, Martin. This may be the Final Solution, but we still may not know the answer!

I reckon you have shown that more than luck is at play to cause differences between armies and their efficacy. Seems all armies are not equal. On balance, efficacy does not tend toward 50% as sample size increases (and as one might expect in balanced army lists) but tends toward a range band of efficacy around 50% with some armies doing better and some doing worse. What makes 127 particularly good and 60 particularly bad?

Vexillia said...

Neat summary as always.

As to 127 & 60: if only I knew. The data doesn't, and can't, tease out any causes.

Personally, I suspect the Classical Greeks (60) are not good outside their period (too slow).

As to 127 (Nikephorians) I wish I knew where the magic came from. There's every chance it's a lot to do with player skill and not just the list. The rules change from V3 to V4 didn't really affect them directly. They just became more popular and performed better.

The analysis is definitive but frustrating as it's a dead end.

Vexillia said...

Prompted by Jon's comment I looked at how the Nikephorian list changed from V3 (128) to V4 (127):

- V4 switches to heavy cavalry impact/½ bow (was impact bow) for Tagmatic cavalry.
- V4 switches to heavy & medium cavalry impact/½ bow (was impact) for Thematic cavalry.
- V4 allows both for the full date range of the list.
- V3 offers Tagmatic & Thematic cavalry but not at the same date (splits at 1042).

So quite a big change in the list as players can field large numbers of impact/½ bow cavalry. It's got to be a factor. Sadly, the analysis can't confirm this.

Salute The Flag

If you'd like to support this blog why not leave a comment, or buy me a beer.