Is Backgammon Luck or Skill?
42,000 measured matches between engines of known strength — and the curve nobody publishes
Last updated
Ask this question anywhere and you get the same answer: backgammon is about 80% luck and 20% skill. Nobody says where the number comes from, what it counts, or over how long a contest it is supposed to hold — and that last omission is the one that matters, because the same two players produce wildly different results over one game and over a long match. So we measured it. Two engines, identical in every respect except how well they play their checkers, across 42,000 matches at seven different lengths. The short version: one game is very nearly a coin flip, and a long match is not. Every figure on this page comes from a tool you can run yourself.
The short answer
Backgammon is a game of skill played through a thick layer of noise. Both halves of that sentence are load bearing, and the balance between them is decided by one thing: how many games you play.
In our measurements, a clearly better player — one who wins 55 of every 100 single games — wins about 70 of every 100 matches to 25 points. Nothing about the players changed between those two numbers. Only the length of the contest did.
That is why the usual question is badly posed. There is no single luck-to-skill ratio for backgammon, any more than there is a single ratio for a coin-flipping contest: flip once and it is all luck, flip ten thousand times and a coin biased by 1% wins nearly every time. What backgammon has is a curve, and the curve is measurable.
Why "80% luck" is not an answer
The 80/20 figure appears on dozens of sites, always without a citation. It is not obviously wrong — a single game really is dominated by dice — but as stated it cannot be true in general, because it does not mention the length of the contest, and the length changes the answer more than anything else does.
The most rigorous published work we could find takes a different approach and reaches a compatible conclusion. A 2016 statistical analysis of roughly 650 real matches, using XG's per-decision equity analysis, found that predicting match results from luck alone was 97.9% accurate while predicting from skill difference alone was 65.0% accurate. That sounds like a knockout for luck until you notice what it measures: within the matches actually played, on the day, the dice explain the result. It says nothing about what happens when you keep playing. The author notes in passing that "the longer the match, the more chance for skill to emerge as dominant" — and then, like everyone else, does not quantify it.
That gap is the reason this page exists. The qualitative claim is repeated everywhere; the curve behind it does not appear to have been published. So we produced it.
How we measured it
The difficulty with measuring skill is finding two players whose difference you can hold fixed and know exactly. Humans will not do: they improve, tilt, tire, and have good days. Two copies of an engine will.
Both sides here are the same engine that plays you on this site, with identical evaluation, identical search and an identical doubling policy. The only difference is that the weaker copy sometimes plays the second- or third-best turn instead of the best one. That is a real difference in checker play and nothing else, and it has a dial: raise the rate and the gap widens. We calibrated three settings against 15,000 single games to find out what each was worth, then played matches at seven lengths — 1, 3, 5, 7, 11, 17 and 25 points, the last being the length of a world championship final.
Two controls run throughout, because a measurement with no known quantity in it is worth nothing:
- An equal pairing at every length. Two identical engines must land on 50%, and they do — 50.0% over 3,000 single games, and between 48.5% and 51.2% at every match length. If the machinery had a bias, this line would show it.
- Swapped seats. Every other match puts the stronger engine on the other side of the board, so moving first cannot flatter either player.
All figures below carry a 95% confidence interval. The single-game column is measured over 3,000 games (±1.8 points), each match cell over 1,200 matches (±2.8 points at most).
One game is almost a coin flip
Here is the first result, and for most people it is the surprising one. Over a single game, the three skill gaps produced these win rates:
| Skill gap | Wins per 100 single games | Implied rating gap |
|---|---|---|
| Equal players (control) | 50.0% ±1.8 | 0 |
| Slightly better | 53.8% ±1.8 | ≈26 points |
| Clearly better | 55.4% ±1.8 | ≈38 points |
| Much better | 62.0% ±1.7 | ≈85 points |
Read the bottom row again. Our widest gap — a player making materially worse checker plays several times a game — still loses nearly two games in five. There is no setting of the dial at which one side simply wins. That is not a weakness of the experiment; it is the nature of the game, and it is exactly why a beginner can sit down against a champion, win, and leave convinced the whole thing is a dice game.
The rating column converts each win rate into the kind of number our leaderboard uses, by the standard Elo relation between expected score and rating difference. It is a translation of the same measurement, not a separate result.
What a match does to that
Now play the same two engines for points instead of for a single game. Nothing about how they play changes. Only the finish line moves.
| Match length | Slightly better | Clearly better | Much better | Control |
|---|---|---|---|---|
| 1 point (single game) | 53.8% | 55.4% | 62.0% | 50.0% |
| 3 points | 53.9% | 57.3% | 63.3% | 48.5% |
| 5 points | 53.0% | 59.8% | 65.3% | 50.4% |
| 7 points | 55.2% | 60.7% | 67.3% | 48.5% |
| 11 points | 52.8% | 63.0% | 71.4% | 49.9% |
| 17 points | 58.1% | 64.2% | 75.1% | 50.5% |
| 25 points | 59.6% | 69.7% | 79.7% | 49.5% |
The clearly better player gains 14 percentage points — from 55.4% to 69.7% — purely by playing to 25 instead of to 1. The much better player goes from 62% to nearly 80%. The control does not move at all, which is what tells you the movement in the other three columns is real.
Why length works: a match is many games
There is no mystery in the mechanism. A match to 25 points is not one long game — it is a series of ordinary games whose results are added up. In our control runs the average match took:
| Match length | Games actually played |
|---|---|
| 1 point | 1.0 |
| 3 points | 1.9 |
| 5 points | 2.8 |
| 7 points | 3.9 |
| 11 points | 5.9 |
| 17 points | 9.1 |
| 25 points | 13.5 |
Fewer games than points, because gammons and the doubling cube mean a single game can be worth two, four or more points. Still, a 25-point match is around thirteen chances for the better player's edge to show, and dice that ruin one game have to ruin most of thirteen to decide the match.
This is the same arithmetic that governs any repeated contest, and it is why serious backgammon has always been played as matches rather than single games. Tournament formats did not choose 7, 11 or 25 points by accident: they are buying skill resolution with time.
The doubling cube makes it luckier, not less
This one we did not expect, and it is worth stating carefully because it is easy to misread.
We ran the middle skill gap twice at every length: once with the doubling cube in play, once with it switched off entirely. Both sides always used the same cube policy, so neither player could out-cube the other. The result:
| Match length | Cube in play | No cube | Difference |
|---|---|---|---|
| 3 points | 57.3% | 59.5% | +2.2 |
| 5 points | 59.8% | 61.3% | +1.4 |
| 7 points | 60.7% | 65.1% | +4.4 |
| 11 points | 63.0% | 68.8% | +5.8 |
| 17 points | 64.2% | 72.0% | +7.8 |
| 25 points | 69.7% | 76.9% | +7.3 |
With cube skill held equal, the cube costs the better player about seven points of match win rate over 25. The reason is straightforward: the cube multiplies stakes, so a match reaches its target in fewer games, and fewer games means less room for an edge to assert itself. The cube is, in this narrow sense, a variance amplifier.
What this does not show is that the cube reduces skill in real backgammon. It does the opposite: cube judgement is one of the largest sources of expert advantage there is, and our engine's cube policy is admittedly crude — it doubles late and takes almost everything, which we document on the doubling cube page. By giving both sides the same crude policy we deliberately removed cube skill from the experiment in order to isolate checker skill. The honest reading is: the cube's raw effect on variance favours the underdog; its effect through judgement favours the expert; this measurement only sees the first.
There is a built-in check on this comparison. At one point there is no cube in either column, so both should measure the same thing — and they do: 55.4% against 54.7%, well within the interval.
Luck never disappears
It would be tidy to end with "play long enough and skill wins". The data does not say that.
At 25 points — the longest match played anywhere, reserved for world championship finals — our widest skill gap still lost one match in five. The clearly better player lost almost one in three. There is no match length in practical existence at which the better backgammon player is safe.
That is a feature, not a defect. It is the reason a club night stays interesting, the reason weaker players keep coming back, and the reason backgammon has survived five thousand years while games of pure skill shed everyone who cannot win at them. A game in which the better player always wins is a game most people stop playing.
Backgammon, chess and poker
The comparison people reach for first is chess, and it is genuinely instructive. Chess has no random element whatsoever: between a strong player and a weak one, the stronger wins essentially every game, and a single game is already a reliable verdict. Our widest gap in backgammon produced 62% over one game. Skill in chess is expressed per game; skill in backgammon is expressed per match.
Poker sits on the same side of the line as backgammon, and poker players will recognise every argument on this page — the distinction between decisions and results, the insistence on a large sample before judging anyone, the fact that a bad player can win a session and a good player can lose for months. Backgammon is simply a much faster version of the same statistics, with the noise concentrated in two dice rather than a shuffled deck.
There is one respect in which backgammon is more resolvable than either: it is solved to the point that a computer can tell you the equity of any position. That is what makes measurements like this one possible at all, and it is why modern backgammon skill is measured by error rate per decision rather than by results.
What a court decided
The question is not only academic. In State of Oregon v. Barr (1982) the state prosecuted tournament director Ted Barr for promoting gambling, arguing that dice made backgammon a game of chance and that prize money therefore made it gambling. The defence called Paul Magriel, author of the standard textbook, who testified that "it is the personal move, the decision where to move your man after the dice have been cast, that is the essence of the game. The roll of the dice does not force your play; it merely reduces your options."
Judge Stephen S. Walker concluded that backgammon is a game of skill rather than a game of chance and acquitted Barr. Gambling law differs from country to country and this page is not legal advice — but the reasoning matches what the numbers show. The dice supply the options; the player chooses among them, and over enough choices the choosing decides.
What this means at the board
Four practical consequences follow directly from the curve.
- Play longer matches if you want the better player to win — including when that is not you. A single game tells you almost nothing. If you are the underdog, short matches are your friend; if you are favourite, they are how you get upset.
- Judge your decisions, not your results. Over ten games the result is mostly dice, so using it to measure your play is measuring noise. This is the single most valuable habit the numbers support, and the hardest one to keep.
- A losing streak is not evidence of anything. If a 62%-favourite loses two games in five, losing four in a row happens roughly once every eighty games by chance alone. On this site the dice are rolled by the server with a cryptographic generator and sent to both players simultaneously, and they will still do this to you, because that is what real randomness looks like.
- Small edges are worth having. The difference between our "slightly better" and "clearly better" settings is under two percentage points per game — and 10 points of match win rate at 25. Learning the shot odds or fixing one habitual error is worth more than it feels like it should be.
What this measurement is not
Stated plainly, because a number without its limits is worth less than no number:
- These are our engine's rates, not backgammon's constants. Our engine is a heuristic evaluator, not a world-class neural bot. What generalises beyond it is the shape — how a given per-game edge grows with match length — because that relationship is statistics, not backgammon. The per-game rate is measured and stated alongside every curve precisely so the shape can be read against it.
- The weaker player is not a human beginner. Ours misplays checkers and nothing else. A real beginner also misplays the cube, misjudges races, and plays worse when behind. A human skill gap is therefore wider than the same per-game rate suggests here.
- Cube skill is excluded by design, as described above. This is the largest single limitation of the experiment.
- The smallest gap is near the resolution limit. At roughly 54% per game, 1,200 matches per cell gives ±2.8 points, so that line wobbles. Its endpoints differ by more than the interval; its individual points should not be read closely.
Where these numbers come from
The measurement is a committed tool in this project rather than a script that was run once and thrown away, because a number nobody can reproduce is not evidence. It writes a single data file, and both this page and the charts above read that same file, so a re-run moves the text and the pictures together.
Totals behind this page: 42,000 matches across seven lengths and five pairings, plus 15,000 further single games to pin down the per-game column, with seats swapped throughout and an equal-strength control at every length.
Published work referenced above:
- Peter Ellis, "Skill v luck in determining backgammon winners" (2016) — a logistic-regression analysis of roughly 650 real matches using XG per-decision data, the source of the 97.9% / 65.0% figures.
- Phil Simborg, "Luck vs. Skill in Backgammon" — the standard essay on the question, argued from professional experience rather than data.
- "The Trial (and Tribulations) of Oregon Promoter Ted Barr", Backgammon Times — contemporaneous account of State of Oregon v. Barr, including Magriel's testimony.
Quick reference
| Question | Measured answer |
|---|---|
| Does a better player win a single game? | Usually not by much — 54% to 62% depending on the gap |
| Does a better player win a long match? | Yes — 60% to 80% at 25 points, same players |
| Do equal players stay at 50%? | Yes, at every length (48.5%–51.2%) |
| How many games is a 25-point match? | About 13 |
| Does the cube help the better player? | Not through variance — it costs about 7 points at 25 |
| Can the weaker player win a 25-point match? | Yes, about one time in five |
| Is one game enough to judge anybody? | No, and neither are ten |
Questions people ask
Is backgammon a game of luck or skill?
Both, and which one decides the result depends almost entirely on how long you play. In our measurements a clearly better player wins about 55 of 100 single games but about 70 of 100 matches to 25 points. One game is close to a coin flip. A long match is not.
How much of backgammon is luck?
There is no single number, which is why the popular 80 to 20 figure is unsatisfying. Luck decides most single games and almost none of a long match. The honest answer is a curve rather than a ratio, and the curve is what this page measures.
Can a beginner beat an expert at backgammon?
Easily, over one game. Even our widest skill gap only produced a 62 percent win rate in single games, so the weaker side still won nearly two games in five. That is the reason backgammon stays enjoyable against stronger opposition in a way chess does not.
Does the better player always win in backgammon?
No, and not even over the longest match anybody plays. At 25 points, the length of a world championship final, the weaker player in our widest gap still won one match in five. Length narrows luck; it never removes it.
Why are backgammon matches played to several points?
Because a single game is too short to identify the better player. A match to 25 points takes about 13 games in our runs, and it is that repetition, not any change to the rules, that lets skill decide the outcome.
Is backgammon more luck than chess?
Yes, unambiguously. Chess has no random element at all, so a stronger player wins essentially every game against a much weaker one. Backgammon has dice, so the same gap in skill produces a 62 percent win rate over a single game in our measurements.
Is backgammon legally a game of skill?
In at least one well-known case, yes. In State of Oregon v. Barr in 1982 the court concluded that backgammon is a game of skill rather than a game of chance, and acquitted tournament director Ted Barr of promoting gambling. Gambling law varies by country and this is not legal advice.
Does the doubling cube add luck or skill?
Both, and our measurement can only show one side of it. With an identical cube policy on both sides, turning the cube on reduced the better player's win rate at 25 points from 77 percent to 70 percent, because the cube multiplies stakes and so amplifies swings. In real play cube judgement is itself a skill, which pushes the other way.
How long does a backgammon match have to be for skill to decide it?
Longer than most people play. In our runs a clearly better player was still under 61 percent at 7 points and only reached about 70 percent at 25. If you want a result that reflects skill rather than dice, play the longest match you have patience for.
What is the 80 to 20 rule about backgammon luck?
It is a widely repeated claim that a single game is roughly 80 percent luck and 20 percent skill. It is quoted without a source and without saying what it counts, and it cannot be right for every contest length, because the same two players produce very different results over one game and over a long match.
Are online backgammon dice fair?
On this site every die is rolled by the server with a cryptographic random generator and sent to both players at the same moment, so no browser and neither player can influence or predict a roll. Losing streaks feel engineered precisely because real randomness is streaky.
How do I improve at backgammon if luck matters so much?
Judge your decisions rather than your results, because over a handful of games the result is mostly dice. Learn the shot odds, learn when to double and when to take, and play longer matches, where better decisions actually show up in the score.
Next: the odds guide gives you the shot numbers that turn into that per-game edge, the doubling cube guide covers the decision this page deliberately held constant, and strategy turns both into plans. If you would rather find out where you sit on the curve, play a rated match — matches to 3 and 5 points are one click from the board.