KO.LOL League of Legends Stats ko.lol Download

How ko.lol computes every number

Six traps decide what a stats site tells you, and this page walks all six

This page is the argument for the whole site. Six statistical traps decide what a League stats tool tells you, and no tool in this market fixes all six. Each one below says what the trap is, what it makes other tools do, what we run instead with the real constants, and one thing you can open right now to check us.

patch 16.14 pool EUW, Korea and NA recomputed

Games recorded
108,616 ranked games in the store
Item events
47,183,210 purchases with the game state they happened in
Rune pages
1,085,538 whole pages counted as one thing
Pairs scored
484 champion and role pairs on patch 16.14
Under the floor
308 pairs under 30 games, no number printed
Pull strength
2,000 pseudo games of the role average

The mechanism for each one is running code with a constant you can read out of our own API.

Before any of it: our numbers come from 108,616 recorded ranked games on EUW, Korea and NA, found by following Diamond players and the people they get matched with. That is a high rank pool on three regions. It is not every rank and it is not every region, and nothing below this line is worth reading if you forget that.

On patch 16.14 that store scores 484 champion and role pairs and refuses to score 308 of them, because a pair needs 30 recorded games before we will put a number on it.

These numbers come from 108,616 ranked games we recorded on EUW, Korea and NA, found by following Diamond players and the people they play with. That is not every rank and it is not every region. Everything on this page is patch 16.14. Counted up to .

The six traps

Each card links the section below that walks the trap in full, with the real constants and one thing to open and check.

Why does the top build on other sites so often flop?

Top of a sorted list
58% the simulation: 50 truly identical builds, 200 games each, and the winner reads this
Real edge found
1 in 3 times the genuinely better option tops the list, against nine equals at 500 games each

The trap

Sort a long list of noisy win rates and the top of the list is not the best option, it is the option that got the most luck. Sorting picks up the largest error along with the largest effect, and it does that even when every option on the list is truly identical.

Simulate it and watch it happen: 50 build options, every one of them truly winning half its games, 200 games behind each. The top of the sorted list reads about 58%. Nothing about that build is real. With one genuinely better option among nine equal ones at 500 games each, the top of the sorted list is the genuinely better one about a third of the time.

What other tools do, and what we do instead

What other tools do

Publish the sorted list raw. Lolalytics sorts matchup cells of a few hundred to a few thousand games into best counters, and has displayed a 595 game item option at 67% win rate with no warning on it. u.gg is the one aggregator that publishes its decision rule, and the rule crowns the most played build that clears a raw win rate floor, so a mediocre build with a large pick rate stays on top of a better rare one.

Fair credit where it is due: DraftGap pulls its small pair samples toward a prior before it uses them, which is hygiene most larger tools skip.

What we do instead

Pull every rate toward the pool it belongs to before anything is ranked. On the tier list the pull is worth 2,000 pseudo games of the role average, so a pair with 30 games sits almost on the role average and a pair with twenty thousand games barely moves. Two lucky weeks never crown anything.

The same shape runs on rune pages with a weaker pull, because a page pool is smaller than a role pool.

Check it: open https://api.ko.lol/v1/tierlist?patch=16.14 and read the formula block at the end of the response. The pull strength, the games floor and the tier cutoffs are in there. The tier list page reads them from that same response, so it cannot print a constant we have quietly changed in the page.

How many games does a win rate need?

One option
±2 pts how far a win rate over 500 games moves on luck alone
Two options compared
±6 pts how far the gap between two 500 game options moves across its 95% range
To split 51 from 52
39,000 games per option, roughly, to tell them apart reliably

The trap

More than a champion and role slice usually has. A win rate measured over 500 games moves about two points either way on luck alone. The gap between two such options moves about six points across its 95% range. Telling a 51% option from a 52% one apart, reliably, takes roughly 39,000 games per option.

What other tools do, and what we do instead

What other tools do

Rank on a tenth of a point. A table that puts one option above another on that margin, over a few hundred games, is showing you rounding and calling it an order. Porofessor prints character labels on 10 to 40 recent games, a sample where a win rate swings about eleven points on luck alone.

Fair credit: Lolalytics prints the sample size on most of its tables, which is the one thing that lets a numerate reader do this arithmetic for themselves, and League of Graphs describes what it counted without dressing it up as advice.

What we do instead

Every estimate prints its 95% range, the games behind it, and one of five words for where it came from. Measured means counted in real games in exactly this spot. Pooled means we borrowed from similar spots and said so. Modeled means we worked it out rather than counted it.

Digits track the range, so a wide range earns fewer decimal places. Under the floor we print how many games there are and no number at all.

Check it: take any estimate on this site, meaning any figure carrying one of those five words, and look for its range and its game count. Every one of them has both, because they are all printed by one component with no way to leave a piece out. The build then reads the pages it just made and stops the deploy if an estimate is missing its range, its games or its word, or if a percentage carrying a decimal point sits anywhere on a page with no game count attached to it.

Where that check stops, so you know what you are looking at: the round numbers in the worked simulations on this page. The 58% at the top of a sorted list of identical builds and the 92% on an item whose real effect is zero are outputs of the simulations described around them. Nothing was counted, so there are no games to print. Every figure on this site that came out of our store is an estimate, and it carries its evidence.

Why do item win rates point the wrong way?

Real effect zero
92% win rate a do-nothing item still shows as a sixth completed item, in the simulation
Sign flipped
20 pts in the wrong direction for a defensive item with a real 5 point benefit, bought mostly when behind

The trap

Items get bought because of the state of the game, and the state of the game decides who wins. So an item win rate counted at the end of the game measures shopping, not the item. Finishing your sixth item is a symptom of winning.

Simulate it: let a hidden advantage drive both winning and item completion, and set every item's real effect to exactly zero. The item still shows a 92% win rate as a sixth completed item. Run it the other way and a defensive item with a real 5 point benefit, bought mostly when behind, shows up 20 points in the wrong direction. The bias is not small and it is not a rounding problem: the sign flips. That is why every site's item tables bury the defensive purchases that a losing game actually needs.

What other tools do, and what we do instead

What other tools do

Count items at the end of the game and sort by that number. "59% of Jinxes finish Berserker's Greaves" tells you who lasted long enough to finish boots. One tool states that it corrects for this but does not publish the correction, and a correction nobody can check is worse than none, because it puts a clean label on a biased number.

Fair credit: Coachless builds its headline metric on the state at the moment of purchase, which is the right conditional and the only one in the market.

What we do instead

Estimate what buying the item changes at the moment you buy it. Buyers are compared with a neutral background stream inside strata of the game state: ahead or behind, crossed with gold difference buckets at 2,000 and 500 gold behind and 500 and 1,500 gold ahead. Two models back each other up, one for who buys and one for who wins.

Buying chances are held between one in ten and nine in ten so a single thin stratum cannot run the whole estimate. The background stream is the long tail of items with fewer than 50 purchase events each, which carry no selection story of their own. The estimate runs in the desktop app, because it needs your live game; how we compute item advice says which half is which.

Check it: when Riot changes an item we also measure it either side of the patch boundary and publish both numbers. Patch 16.14 has 6 of those reports at https://api.ko.lol/v1/rdd?patch=16.14, each with its range and its games on each side.

Why is best of each row not a build?

The real pair
54% what the truly best rune pair wins in the simulation
Each half reads
50% in the per row tables, because the pair is rare and keeps bad company otherwise
Who runs it
4% of players in the simulation, which is why the per row tables miss it

The trap

Rune tables show a win rate per row: one table for keystones, one for the second row, one for shards. Take the best cell in each table and you have assembled a page out of parts that were measured apart from each other. Runes interact, so that page can be one nobody plays and nobody wins with.

Simulate two rune rows with a real interaction. The truly best pair wins about 54% of its games, while each half of it reads about 50% in the per row tables, because only 4% of players run the pair and the rest of the time those choices keep bad company. The per row tables say the winning pair is losing. That is how a real off meta page stays invisible to any tool whose rune tables are built one row at a time.

What other tools do, and what we do instead

What other tools do

Publish per row tables, and in the import apps, assemble the per row winners into a loadout with one click. That assembly is the failure, executed at speed.

What we do instead

Rank whole measured pages. Keystone, both trees, every row and all three shards come from one page's real games, counted together, and the page is ranked as one thing. We never stitch a keystone onto another page's rows. That is true of the rune section on every champion page here and of the rune card in the app.

One part of it is the app's alone. Swapping to a different whole page when your lane opponent is revealed needs to know who you are up against, which the app reads from the draft and a web page cannot. The swap fires only when both pages have real games in that exact matchup bucket, and if no measured page fits, the card says so instead of building one.

Check it: open https://api.ko.lol/v1/runes/Ahri/mid?patch=16.14. Every entry in the pages list is a whole page with its own games and wins. Where a champion page does show one slot on its own, the note under it says that it is one slot counted across every page that used it, and that it is not a build. The whole pages sit above it, and they are what we rank.

Why do draft tools overstate their edge?

The sum claims
92% when the truth is about 60%, in the simulation of five noisy readings of one thing
Pairwise sum
65% for the all magic damage five stack, DraftGap's published worked failure
Whole draft model
40% for the same five stack, because no single pair is the all magic damage problem

The trap

Draft tools built on pairwise statistics add up one number per matchup and one per pairing. Adding them assumes each pair is a separate piece of evidence. They are not. Three enemy champions with heavy attack damage are one reason to build armor, counted three times.

Simulate five signals that are all noisy readings of one underlying thing. Each one measures correctly on its own. Added up the way a pairwise engine adds them, a truth worth about 60% becomes a claim of 92%. The combining rule is what manufactures that certainty, not any one of the numbers going into it. DraftGap publishes its own worked failure: a full magic damage five stack scores about 65% by the pairwise sum and about 40% under a whole draft model, because no single pair of champions is the all magic damage problem.

What other tools do, and what we do instead

What other tools do

Add the pair numbers. Fair credit, and it matters here: DraftGap's formula is open source, which is the only reason anyone can name its failure this precisely. A tool that publishes a wrong formula is more useful than one that publishes nothing.

What we do instead

One model reads the whole seat in one pass. Five small networks vote, all reading the same 54 numbers that describe both teams, and the spread between them sets the range. There are no pair terms anywhere in the output, so there is nothing to double count.

Check it: the draft read runs in the app, because it needs the draft as the client fills it in, and there is no draft screen on this website to look at. In the app it prints one estimate with a range and the word Modeled on it. If you want the other object to compare it against, DraftGap's summing formula is on GitHub in full.

How much can a draft model actually know?

Perfect model
52% how often it calls the winner when the draft moves the true chance by a couple of points
Best published result
55% in the peer reviewed literature, on 279,893 games from the top tenth of a percent
One tool claimed
62% before an outside tester measured about 52% against its public API

The trap

Less than the marketing says. In soloqueue the draft is a small term in a result decided by ten players and forty minutes of execution. If the draft moves the true chance of winning by a couple of points, then a model that knew the true chance exactly would still only call the winner about 52% of the time. That is the ceiling for a perfect model, not for a good one.

The published research lands where that arithmetic says it must. The best result in the peer reviewed literature reaches about 55% on 279,893 games from the top tenth of a percent of players, and the methods before it sit between 52% and 54%. So a tool claiming 58% accuracy is claiming drafts routinely decide games at odds nobody observes.

What other tools do, and what we do instead

What happened in this market

One tool claimed 62%, an outside tester measured about 52% against its public API, and the corrected model settled near 55%. The correction is the honest part of that story and it is why the tool in question has the best method in the market.

There is a subtler point underneath it. Most of the accuracy a draft model earns comes from a tail of broken drafts that are easy to call. In the band where sensible pick decisions actually live, the same models resolve almost nothing.

What we do instead

Publish no accuracy headline at all. The draft read is labeled Modeled with a sample size of zero, because it is a model and not a count of games. Its range widens for every seat where we do not know who is playing.

When the top picks' ranges overlap, the screen says any of these works and shows them as equals, and bans are the biggest threat rather than the most likely.

Check it: hold any accuracy claim against the numbers in the paragraph above. Then look at ours: there is no number to hold, because we do not publish one.

What ko.lol never claims

The rule that comes before every other rule on this page.

  • Rune page win rates are associations, not proof that a page makes you win. Players who take a page into a matchup may be better at that matchup already. No rune tool in this market separates those, us included.
  • The draft read is Modeled with a sample size of zero. It trained offline on drafts we generated, not on a count of real games, and it says so on screen every time.
  • Our store is high rank ranked solo games from EUW, Korea and NA. Low rank cuts, off meta cells and other queues are thin here, and thin cells print their game count instead of a number.
  • Data Riot does not publish gets no stand in from us. Ping timing, true duo flags, per fight detail and full ladder history are not in the match data, so we say that rather than estimating them.
  • Under a floor we show nothing and say how many games there are so far. On patch 16.14 that is 308 champion and role pairs with no score.

What we compute, per recommendation

One page per mechanism. Anything that reads your live game runs in the app, because a static page cannot see one; what this website has is the measured store both of them read.

TL

The tier list

Win impact, presence and mastery, pulled toward the role pool before ranking, with floors and a range on every row. The other way is one sorted table for everyone, no range on any row.

CH

Champion picks

Measured impact with who is playing in the score, and ties shown instead of an invented winner. Most tools publish one table for everyone.

RU

Runes

Whole measured pages ranked as one thing, with ties shown as ties. The other way is per row winners assembled into a page nobody played.

IT

Items

Estimated at the moment of the buy, inside strata of the game state. The other way is counting items at the end of the game, where the sign can flip.

DR

Draft

One model per seat, labeled Modeled, and no accuracy headline. The other way is pair deltas added up, with one.

SC

Your game, graded

Percentiles against recorded games for that champion and role, the pool named, no grade below the floor. The other way is one absolute grade, inputs and weights undisclosed.

Not written yet: the career score page, the jungle route notes and the skill order notes. The mechanisms exist in the app and the server, the pages explaining them do not, and naming them here is more useful than a link that goes nowhere.

How to check any of this

Four checks. None of them require trusting us.

  1. Read our own constants out of our own API. The tier list ships its formula block in every response, so the pull strength, the games floor and the tier cutoffs are one request away: https://api.ko.lol/v1/tierlist?patch=16.14.

  2. Recompute a pull. Take a rune page's wins and games from https://api.ko.lol/v1/runes/Ahri/mid?patch=16.14 and the published pull strength, and you can work out how far we moved it before ranking it.

  3. Compare the same entity on a competitor's page. Same champion, same role, same patch. Count what is printed next to their number, then count what is printed next to ours.

  4. Check us against something outside all of this. The draft ceiling arithmetic and the published research are not ours, and DraftGap's summing formula is open source.

The fine print: our pool, our updates, our corrections

The pool is stated on every page that prints a number, generated from the crawler's own state rather than typed in, so the day a fourth region starts being recorded the sentence changes itself. Right now it reads EUW, Korea and NA.

Stats pages rebuild nightly and again on patch day. Every page stamps the date its numbers were counted up to, in the footer, as a real date a machine can read. Patch reports keep updating for the first days of a patch while the games arrive, then settle.

When we get something wrong we log it rather than quietly editing the page. The log is on the changelog, and it is also where a change to any constant on this page will appear on the day it changes.

ko.lol for desktop

This page answers the question you searched for. The app answers the one you have while the game is running: it reads the draft as it fills in and names the pick, reads your gold and items and names the buy, from the same store of recorded games.

Free. For Windows.

Download