I do. Personally, with my own money, from my own wallet.
My name is George V., though most people in this corner of the industry know me as efialtis. I have been testing crypto bookmakers since 2020, first at BTCGOSU and now here. AffPapa iGaming Awards — Crypto Affiliate of the Year, 2022 and 2023.
That matters for one reason only: when a payout goes wrong, I know what normal looks like. A book that suddenly asks for documents it never asked for, a withdrawal that takes four hours where it used to take six minutes — those are the signals worth publishing, and you only see them if you have been watching for years.
Everything on this site is measured by one person. There is no team, no outsourced testing, no data bought from anywhere. If a figure is here, I put it there.
1. A real account, funded with my own coins. No press account, no arrangement with the operator, no pre-loaded balance. Sometimes the account is new. More often it is one I have held for years — both are marked in the record, because a fresh account goes through first-time verification and an established one does not.
2. Real bets. At least one, sometimes ten. Mostly overs. The point is not the bet, it is that the balance has moved before it leaves — an account that deposits and immediately withdraws triggers automated fraud checks, and then you are measuring the fraud system, not the payout.
3. A withdrawal, without warning. No support ticket first, no notice, nothing that would let anyone treat the request differently. The clock starts at the moment I press withdraw and stops when the coins are spendable in my wallet.
4. Everything published. Transaction ID for the deposit and the withdrawal, the exact timings, the network fee, whether documents were demanded and at what amount. You can check the chain yourself.
The amounts vary — three figures on one test, four on the next. So does the number of bets, and so does the time between deposit and withdrawal.
That is not sloppiness. Varying the amount is the only way to find where the thresholds sit. A book that pays 200 € without a question may demand a passport at 2,000. If every test used the same figure, that line would never show up.
Every record shows its own amount, bet count and holding time, so nothing is hidden behind an average.
| Time to wallet | From pressing withdraw to spendable coins, including network confirmation |
| ID threshold | The amount at which documents were actually demanded — not the stated policy |
| Verification time | How long the check took, and whether the balance was frozen during it |
| Fee behaviour | Whether the house absorbed the network fee, passed it on, or took an extra cut. Calculated from the chain, not from what the cashier claimed |
| Reliability | How many tests ended paid, partially paid, refused, or blocked |
| Recency | When the book was last measured |
Confirmation times, amounts and network fees are read directly from the blockchain and marked as such in each record. Anything typed by hand is marked as typed by hand.
One number in every test has no chain record: the moment I pressed withdraw.
That click happens inside the operator’s interface and leaves no public trace. Everything after it is verifiable — the confirmation, the amount, the fee. The starting point is not.
The honest answer is a screenshot of the withdrawal request with the time visible, filed with every completed test. It is not proof in the way a transaction ID is proof. It is the best that exists, and pretending otherwise would undo the point of this site.
Every book is retested on a fixed interval — quarterly by default, and monthly for the ones I decide need watching more closely. That is a judgement, not a measurement. Nothing on this site counts how often readers ask about a book, so this page does not claim it does: the interval is recorded against each bookmaker, and it is mine to set.
A figure that has not been remeasured in 90 days is flagged stale on its page. Not quietly left standing — flagged, so you can see that the number is old.
Books get worse as well as better. A payout time from March says nothing about August.
| After | What happens |
|---|---|
| 48 hours | The test is marked overdue and shows as overdue while it stays open |
| 7 days | I contact support, document the exchange, and file it as an incident on the book’s record |
| 30 days | The test is recorded as refused |
Support is contacted only after the test has already gone overdue. Contacting them earlier would turn an unannounced withdrawal into an announced one, and the measurement would be worthless.
Scores are calculated, not awarded. Every input has a fixed weight, and no operator sees a score before it is published.
| Input | Weight | Source |
|---|---|---|
| Payout speed | 30 % | Median across all tests |
| ID behaviour | 20 % | Level of ID demanded, how long the check took, whether the balance was locked |
| Reliability | 15 % | Share of tests ending paid |
| Licence | 15 % | Jurisdiction, plus whether the number was verified in the register |
| Documented incidents | 10 % | Date, type, source and severity of each entry |
| Bonus value | 10 % | Calculated from wagering requirement, basis, game weighting, max bet and time limit |
The weights above say how much each input counts. This section says how each one is arrived at, in enough detail that you can take the raw figures out of the payout ledger and reach the same number I did. A score that cannot be recalculated by a reader is an opinion with a decimal point.
This is meant to be checked. Every measurement behind the scores is published as JSON at /api/payout-tests.json — one row per test, carrying the timings, the outcome, the amounts, the fees and the course of any identity check. Where a figure is derived, the figures it was derived from are published beside it: the duration of an ID check sits next to the moment documents were demanded and the moment they were accepted, so the derived number can be recalculated rather than taken on trust. Take the rows for a single book, apply the rules below, and you should arrive at exactly the payout speed, ID behaviour and reliability scores shown on its page. If you do not, one of us has it wrong, and I would rather hear about it than not.
The other three inputs are not in that file, and it is worth saying so plainly. Licence, incidents and bonus are not measurements at all — they are recorded facts about the book, published on its own page with the source attached. Check those against the register, the source link and the operator’s terms rather than against this file. They will get an export of their own once there is something in them worth exporting.
Every scale works the same way. A handful of anchor points fix the score at particular raw values, and anything between two anchors is interpolated. Anchors rather than a formula, because each one can be argued about on its own: “48 hours is zero” is a claim you can dispute. An exponent is not.
Below the lowest anchor the score stays at the top value; above the highest it stays at the bottom one. Those are deliberate floors and ceilings, not extrapolation.
Between two anchors (x₁, y₁) and (x₂, y₂), a raw value v scores:
y₁ + (y₂ − y₁) × (v − x₁) / (x₂ − x₁)
— except for the two time-based scales, where the same expression is applied to the natural logarithms of the values instead:
y₁ + (y₂ − y₁) × (ln v − ln x₁) / (ln x₂ − ln x₁)
Why logarithmic for time. A payout scale runs from five minutes to two days — nearly three orders of magnitude. Interpolated linearly, the midpoint between five minutes and 48 hours falls at 24 hours, which would score 5 out of 10: a scale on which a full day of waiting is “average”. Waiting is not experienced that way. Thirty minutes instead of five is noticed; six and a half hours instead of six is not. On a logarithmic scale each equal step is an equal multiple, which matches how the delay actually feels. The bonus scale is linear, because required turnover is an amount you have to put through, and twice the figure is twice the work.
A test contributes a duration in two cases: it finished, in which case the duration is the measured one; or it is still open and has already passed its deadline, in which case the time elapsed so far is used. A test still running inside its window contributes nothing — it is neither fast nor slow yet, it is unfinished.
The overdue rule is not a detail. Without it a stalled payout would have no duration at all and would drop silently out of the median, and a book could improve its payout score by not paying.
The median is taken over those durations. With an even number of values it is the mean of the two middle ones.
| Median wait (crypto) | Score |
|---|---|
| 5 minutes or less | 10 |
| 15 minutes | 8 |
| 1 hour | 6 |
| 6 hours | 4 |
| 24 hours | 2 |
| 48 hours or more | 0 |
The four middle rows are the same class boundaries the payout ledger uses for its distribution chart. Zero sits at 48 hours because that is where a payout is publicly marked overdue on this site: a wait we call overdue cannot also earn points.
Worked example. A median of 37.5 minutes falls between the 15-minute and the 1-hour anchor:
ln 37.5 = 3.6243 ln 15 = 2.7081 ln 60 = 4.0943
(3.6243 − 2.7081) / (4.0943 − 2.7081) = 0.6610
8 + (6 − 8) × 0.6610 = 6.68
You can check the arithmetic against a real measurement today: the one BC.GAME payout in the ledger took 22 minutes, which by the same steps gives 7.45.
Bank transfers are scored on their own scale, in working days, and interpolated linearly — the range spans one order of magnitude, not three. Weekends are excluded; public holidays are not, because they differ by country and a half-corrected figure would be worse than a stated gap.
| Median wait (bank transfer) | Score |
|---|---|
| Same working day | 10 |
| 1 working day | 8 |
| 2 working days | 4 |
| 3 working days or more | 0 |
Where a book has been tested both ways, each method class is scored against its own table and the two are combined in proportion to how many tests fall in each. Three days is unremarkable for a transfer and catastrophic for a crypto payout; averaging them on one scale would compare two different things.
Scored per test, then averaged. What is scored is friction: what was demanded, how long it took, and whether the money stayed reachable.
| Demanded | Level score | Time the check took | Duration score |
|---|---|---|---|
| Nothing | 10 | 1 hour or less | 10 |
| Email confirmation | 8 | 4 hours | 8 |
| ID document | 6 | 24 hours | 6 |
| ID plus proof of address | 4 | 72 hours | 3 |
| 7 days or more | 0 |
Where no check was demanded, the Level score is the test’s score on its own. Where one was demanded, the two figures count equally: their mean. A check still open is scored on the time elapsed so far, for the same reason a stalled payout is — otherwise the longest check would be the one that never counts. A balance locked during the check costs a further 2 points. The result is capped at 0 and 10.
Worked example. A test where an ID document was demanded (6) and the check took two hours — interpolated between the 1-hour and 4-hour anchors as ln 120 against ln 60 and ln 240, giving exactly 9.0 — scores (6 + 9) / 2 = 7.5. If the balance was locked while it ran, 5.5.
Tests run on an account that was already verified before the test began are left out of this average entirely, rather than counted as “nothing demanded”. No first-time check was observed there, and recording that as good behaviour would be a claim the test cannot support.
What is deliberately not scored is the threshold — the amount at which documents get demanded. Every test records that amount, and varying the amounts across tests is the reason it ever becomes visible. But turning it into a score would need all the amounts in one currency, and applying an exchange rate to a past date invents a precision that does not exist: whether 0.3 BTC was 18,000 or 26,000 euros at the time was decided by the day, not by the bookmaker. The threshold stays where it is checkable, on the individual record.
Each concluded test earns a share of one point, and the score is ten times the average.
| Outcome | Counts |
|---|---|
| Paid | 1.0 |
| Paid partially | the share that actually arrived |
| Refused, or blocked by KYC | 0 |
| Still open, past its deadline | 0 |
| Still open, inside its window | not counted either way |
The share for a partial payout is the amount received divided by the amount that should have arrived — that is, the requested sum less any fee the house charges, and less the network fee where the record shows the network fee was passed on. Deducting a chain fee correctly is not a partial refusal, and without that correction it would score as one.
Worked example. Five concluded tests: three paid, one that returned half, one refused. (1 + 1 + 1 + 0.5 + 0) / 5 = 0.7, so 7.0 — before the correction for a thin record described below, which brings it to 6.25.
The jurisdiction sets a base score, on one question: what happens to a player who has been wronged. Is there an independent complaints route, are player funds required to be kept separate, and has the regulator ever actually sanctioned anyone.
| Jurisdiction | Base |
|---|---|
| Malta, Isle of Man, Gibraltar | 10 |
| Kahnawake | 7 |
| Curaçao | 4 |
| Anjouan, Tobique | 2 |
| Costa Rica, or no licence at all | 0 |
Costa Rica scores zero because Costa Rica does not license gambling. What is sold as a “Costa Rica licence” is a company registration: there is no gaming register to check it against and no authority to complain to. That a company exists somewhere is no help to anyone whose balance was confiscated.
The register check then adjusts that base:
| Register check | Effect on the base |
|---|---|
| Found in the register | unchanged |
| Searched — not in the register | score is 0, whatever the jurisdiction |
| Register unreachable | × 0.8 |
| Checked more than 90 days ago, or not dated | × 0.8 |
| Not checked yet | no score at all — see below |
The second row is an override rather than a multiplier on purpose. A claimed Malta licence that is absent from the Maltese register is worse than an honestly held Anjouan one, and any multiplier short of zero would still leave the false claim scoring higher than the honest one.
Worked example. Curaçao, found in the register, checked last week: 4.0. The same licence with a check from a year ago: 4.0 × 0.8 = 3.2.
This input starts at 10 and each incident subtracts:
severity × source × recency × (settled ? 0.5 : 1) × (predecessor ? 0.5 : 1)
| Severity | Deduction | Source | Counts for |
|---|---|---|---|
| Worth noting | 0.5 | Regulator or court | 100 % |
| Minor | 1.5 | Our own test | 100 % |
| Serious | 4 | Statement by the operator | 90 % |
| Disqualifying | 10 | Trade press | 70 % |
| Forum or user report | 30 % | ||
| No severity — chronicle only | none |
An operator counts high because a bookmaker who is himself the source of an incident is making an admission, which is firmer than a press report.
Not every entry in a book’s chronicle is a charge against it. A move from one licensing jurisdiction to another is recorded, because without it the sequence makes no sense: a licence ends in one place and starts in another, and a reader who cannot see the join is left guessing whether something was hidden. But changing jurisdiction is not misconduct. Entries of that kind carry no severity and subtract nothing — they are there to explain the chronology, not to weigh on it. Whatever was damaging about the circumstances is recorded in the entries describing those circumstances, and counted there; counting it again at the move would be counting it twice.
Recency holds at 100 % for the first twelve months, then falls in a straight line. In months m beyond the first twelve the weight is 1 − 0.7 × (m − 12) / 36 — but it never falls below a floor, and the floor depends on how serious the incident was.
| Severity | Never falls below | Reached after |
|---|---|---|
| Worth noting | 20 % | never fully — it keeps decaying to 20 % |
| Minor | 30 % | 4 years |
| Serious | 50 % | about 2 years 2 months |
| Disqualifying | 80 % | about 22 months |
Why the floor is not the same for all of them. A single floor of 30 % produced a contradiction on this page: the heaviest grade is called disqualifying, and after four years it cost three points out of ten — two, if the only source was trade press. A book that had kept player money would then have looked like one with two small blemishes. The label and the arithmetic were saying different things.
The question is not how old an incident is but how long it still says something about the house. A note-worthy oddity from 2019 can fade to almost nothing. A book that took player funds has told you what its worst case looks like, and that does not stop being true because three years passed. So a disqualifying incident still costs eight points out of ten after a decade, while a minor one settles at three.
Nothing ever reaches zero. A settled incident is meant to weigh less, not to vanish, and a weight that decays to nothing is a record that quietly disappears.
Where only the month or the year of an incident is established, the record says so, and the ageing above is calculated from the first day of that period. Across a decay measured in years a few days change nothing — but the distinction is kept, because it is a statement about the evidence and not a matter of formatting. A day that was established and a day entered because the field demanded one are not the same thing, and only a separate note can tell them apart. What you see is what is known: “October 2024”, not “1 October 2024”.
Worked example. One serious incident reported by the trade press, a month old and still open: 4 × 0.7 × 1.0 × 1 = 2.8, so this input scores 10 − 2.8 = 7.2.
When the company behind the brand has changed. An incident sometimes concerns a company that no longer operates the site. Where that company’s liabilities were not taken over by the one running it today, the incident counts half.
It counts half, not nothing, and the record says so on the page rather than quietly shrinking a number. You would still be dealing with the same brand, at the same address, running the same software. Whoever buys a brand buys its reputation.
Four conditions have to be met before that reduction applies, and they exist because a company can be renamed in an afternoon:
| Condition | Why |
|---|---|
| Recorded explicitly as a predecessor whose liabilities were not assumed | The default is “not established”, and that reduces nothing. An open question is a gap, not a finding. |
| The predecessor’s company registration number | A name is an assertion. A number can be looked up. |
| A separate source for the succession itself | That something happened and who answers for it are two questions with two sources. A trade-press summary of the incident does not establish the second. |
| The incident is not disqualifying | See below. |
A disqualifying incident is never reduced, whatever happened to the company. This is the part worth reading twice, because it decides what the whole rule is for.
If losing player funds could be cleared by re-registering under a new name, then re-registering would be the cheapest thing an operator could do — and this page would be the thing that made it worth doing. A scoring rule that rewards restructuring produces restructuring. So the reduction stops exactly where the incentive would start: a book that took player money carries that at full weight for as long as this site exists, under whichever company name it trades.
The rule is built against misuse, not for leniency. It exists so that a genuine change of ownership can be recorded honestly — with a register number and a source — and for no other reason.
An incident with no checkable source does not count at all. A source link does the job here that a transaction ID does in a payout test, and the rule is the same: without one this is a claim, not a record. The exception is our own tests, where the test itself is the evidence.
What is scored is the turnover you have to put through, measured against your own money. The size of the offer does not enter it: a larger bonus on a worse condition is not a better bonus. With a wagering factor w and a match percentage p:
| Wagering applies to | Effective requirement |
|---|---|
| The bonus | w × p/100 |
| The deposit | w |
| Deposit plus bonus | w × (1 + p/100) |
Measuring against the deposit rather than the bonus is what makes the three comparable. A 200 % match at 40× the bonus would otherwise look identical to a 50 % match at 40×, though it demands four times the turnover.
That figure is then divided by the sports weighting, expressed as a fraction: this is a sports betting site, and a bonus clearable only on slots is not a bonus for the people reading. A sports weighting of zero is treated as no bonus at all.
| Effective requirement | Score |
|---|---|
| 10× or less | 10 |
| 20× | 8 |
| 30× | 6 |
| 40× | 4 |
| 60× | 2 |
| 80× or more | 0 |
Two conditions then subtract, because a requirement is only as fair as the time and the stake allowed to meet it. A time limit under 7 days costs 2 points, under 14 costs 1, under 30 costs 0.5. A maximum bet below 5 % of the bonus costs 1.5 points, below 20 % costs 0.5.
Worked example. A 100 % match with 35× on deposit plus bonus, clearable on sports at full weight: 35 × (1 + 1) = 70× effective, which falls between the 60× and 80× anchors at 2 + (0 − 2) × (70 − 60)/(80 − 60) = 1.0. Cap the stake at 5 on a 200 bonus — 2.5 % — and it drops the further 1.5, to 0.
A book that offers no welcome bonus scores 5 here. That is an explicit neutral rather than a measurement: no bonus is neither an advantage nor a trap, while an unclearable one is a trap.
Each input’s score is corrected for the thinness of the record where it rests on tests, multiplied by its weight, and the six are added and divided by 100. Taking the worked examples above as a single book:
| Input | Measured | Tests | Published | Weight | Contributes |
|---|---|---|---|---|---|
| Payout speed | 6.68 | 4 | 5.96 | 30 % | 1.79 |
| ID behaviour | 9.10 | 5 | 7.56 | 20 % | 1.51 |
| Reliability | 7.00 | 5 | 6.25 | 15 % | 0.94 |
| Licence | 4.00 | — | 4.00 | 15 % | 0.60 |
| Documented incidents | 7.20 | — | 7.20 | 10 % | 0.72 |
| Bonus value | 0.00 | — | 0.00 | 10 % | 0.00 |
| Total | 5.6 |
The three inputs that come from tests differ between Measured and Published; the other three do not, because they do not rest on a sample. Licence, incidents and bonus are read off records, not estimated from a handful of observations.
The Contributes figure is the useful one. It shows at a glance that this book loses almost all of the 2.5 points available from its licence and its bonus, while its actual payouts — the thing it is measured on hardest — deliver most of what they can.
A figure calculated from one measurement is not wrong, but it is not yet reliable — and nothing about the figure itself shows that. A single paid test scores ten out of ten on reliability, at every bookmaker, including one that in truth refuses every second withdrawal.
So each of the three inputs that come from tests is pulled towards a fixed anchor, in proportion to how thin the record is:
(n × measured value + 3 × 5.0) / (n + 3)
The pull fades on its own. With one test the measurements carry a quarter of the published figure, at three tests half, at ten more than three quarters, at twenty seven eighths. Nothing has to be switched off once there is enough evidence.
Worked example. The one BC.GAME payout took 22 minutes, which on the scale above is 7.45. With a single measurement behind it:
(1 × 7.45 + 3 × 5.0) / 4 = 5.61
With three tests at the same median it would be 6.22, with five 6.53, with twenty 7.13. The published figure moves towards the measurement, never away from it.
Why the middle of the scale, and not the average of the books we have measured.
An average across the books already tested would be measurably more accurate. We ran it: against an anchor of 6.5 the mean error on a single measurement falls from 1.38 points to 1.07. We do not use it.
It would build in an assumption we cannot support — that a bookmaker nobody has measured behaves like the ones we have. That is precisely the question a test is meant to answer. An anchor that answers it in advance makes the score self-confirming exactly where it rests on the least evidence, and a new book would start out carrying the reputation of its neighbours rather than its own.
So the anchor is 5.0, the middle of the scale: neither good nor bad, because with no measurement both are equally likely. The cost is a little more error, in exchange for an assumption that can be defended without pointing at a statistic that does not exist yet. And it pulls in both directions — a thin record with a poor result is lifted towards the middle by exactly as much as a thin good one is held back. A single payout scoring 2.0 is published as 4.25.
Below a minimum basis there is no rating at all. The page says what is missing instead, in the same way the ledger withholds a median until there are enough measurements for one to mean anything.
| Required before a score is published |
|---|
| Three published payout tests |
| At least two of them with a usable duration |
| Licensing jurisdiction established, and the register actually checked |
| A date on which incidents were searched for — an empty list is worth nothing without one |
| An answer to whether a welcome bonus exists, and its terms if it does |
Three is lower than it once was, and the reason is the correction described above. We simulated it: three tests with the correction applied carry the same average error as five without it — 0.36 points across the three test-based inputs, the same figure to two decimals. The waiting was buying accuracy that the correction now supplies, so it was dropped.
\nThe payout ledger still shows no median below five measurements, and that is not an inconsistency. The ledger publishes a median raw, as a headline figure across every book we have tested; a raw median needs more behind it than a corrected score does. Two different numbers, two different jobs.
If any one of the six inputs cannot be scored, the other five are not scaled up to fill the gap. Reweighting would produce a complete-looking number out of incomplete data, and it would do so most readily on the books where only the flattering fields have been filled in. There is no score until all six are there.
A missing figure is never treated as a bad one. “Not checked” and “checked, nothing found” are opposite findings, and an unchecked field is my omission, not a verdict on the bookmaker.
Two consequences worth stating plainly.
A book can be excellent at one thing and poor at another, and the score will show both. Fast payouts do not cancel out a regulatory investigation. The breakdown is published on every bookmaker’s page, so you can see which input pulled the number where.
The licence figure distinguishes four states, not two: verified in the register, searched but not found, register unavailable, and not yet checked. A tick box cannot tell “I looked and found nothing” apart from “I have not looked” — and that difference is exactly what matters.
I earn affiliate commission when a reader signs up through a link here. That is what funds the testing, and there is no other revenue.
Three things follow from that, and they are not negotiable.
Commission rate is not an input to the score. It is not in the table above and it never will be.
Every book was funded out of my own wallet before anything was written about it. That includes books I earn nothing from — if a bookmaker is worth measuring, it goes in the ledger regardless of whether there is a deal.
When a partner tests badly, the score drops and the link goes. The book stays on the list with the reasons attached — quietly removing a bad result would be worse than never publishing it. But it loses its “Visit” button and keeps only “Review only”. A recommendation link on a book I would not recommend is exactly the conflict this site exists to avoid.
I do not rank by commission. I do not accept payment for placement, and no operator has ever seen a score before publication.
I do not publish a review without a test behind it. If there is no measurement, there is no page.
I do not average away the bad tests. Every result stays in the record, including the ones that make a book I earn from look worse.
And I do not claim to have tested things I have not. That last one is why this site was rebuilt.
Tell me. The Bitcointalk thread is read daily, and corrections get published rather than quietly fixed.
If a bookmaker has treated you badly — or unexpectedly well — send it over. Reader reports are shown separately from my own measurements and never feed into the score, but they do decide what gets tested next.
Last updated: 28 August 2026