What Galley does, what the numbers mean, and what to do when the answer is not the one you wanted.
You give it the results of an A/B test: how many visitors each version got, how many converted, and optionally how much revenue and how many units they produced. It runs the standard statistical tests on those numbers and tells you whether the difference between the versions is real or whether it is the kind of gap you would expect from chance alone.
The answer comes as a sentence at the top of the report, followed by the detail behind it.
No. That is the point of it. You do not choose a test, set a parameter, or interpret a coefficient. Galley picks the right tests based on what you entered and writes the conclusion in plain language.
The technical detail is all on the page if you want it, and every section explains what it is showing and why it matters. But you can act on the verdict without reading any of it.
The minimum is two per variation: visitors and conversions. That is enough to test significance on conversion rate.
Add revenue and units sold and the report also checks revenue per visitor and average order value, which is where hollow wins get caught. Add the duration and start date and it can tell you whether you have run long enough.
Usually not, but the honest answer is that it depends on the cost of being wrong.
"Not significant" means the difference you are seeing is within the range that chance could produce, so you cannot yet tell a real effect from noise. If the change is cheap and low risk, shipping on a weak signal may be fine. If it touches checkout or pricing, it is not.
Look at the confidence interval rather than the headline number. If the plausible range still includes zero, you genuinely do not know which version is better yet. Galley tells you how much longer you would need to run to find out.
Long enough to reach the sample size the effect requires, and never less than one full week.
The week matters because shopping behaviour is not the same on a Tuesday as it is on a Sunday. A test that runs Monday to Thursday has measured weekday shoppers, not your customers. Two full weeks is safer, because it covers the weekly cycle twice and absorbs a one-off spike.
Stopping the moment a result looks significant is the most common way to manufacture a fake winner. Decide the duration before you start, and let it finish.
Because they can move in opposite directions, and when they do, conversion rate is the one that lies to you.
A variation that pushes a cheaper product, or strips away an upsell, can convert more people while making less money per visitor. On a conversion rate dashboard that reads as a clear win. Revenue per visitor is the number that reflects what actually landed in the bank, so Galley reports it every time you supply revenue.
You split traffic fifty-fifty, but the results come back 52 to 48. A small gap is normal. A large one means something went wrong in how visitors were assigned: a redirect that failed, a bot filter that hit one version harder, a tag that fired late.
It matters because it breaks the assumption the whole test rests on, which is that the two groups are otherwise identical. If they are not, no amount of statistical work on the outcome will save you. Galley runs the check automatically and warns you at the top of the report, before you read anything else.
Because testing more ideas at once gives chance more opportunities to produce something that looks like a winner. Run enough comparisons and one will cross the line by luck.
Galley applies a Holm correction, which tightens the threshold in proportion to how many variations you are comparing. It is the reason a result that looked significant on its own may not survive once it is judged alongside its siblings. This is a feature, not a penalty: without it, multi-variation tests routinely produce winners that do not replicate.
Galley recognises the export shapes from VWO and Intelligems directly, along with the common formats used by Optimizely, AB Tasty, Convert, Kameleoon, Adobe Target, GrowthBook and PostHog.
If it does not recognise a file, nothing breaks. It shows you the columns and asks you to point out which is which, then previews the result so you can confirm the numbers before anything is filled in. That path works with any spreadsheet, including one you built by hand.
Your CSV is parsed in your browser and is never uploaded. The file itself does not leave your computer.
Only the summary counts you can see in the form, the visitors and conversions and totals, are sent to run the analysis. Order-level detail, customer information and anything else in the original export stays local. There is no account, so nothing is stored against your name.
No. The verdict, every statistical test, the confidence intervals and all the charts run without any AI involved. That is arithmetic, and it costs nothing.
The AI layer is optional and adds two things: a written report structured for a stakeholder, and a short explanation under each chart. There is also a rule-based report that needs no key at all and makes no external call, if you want prose without connecting anything.
It is saved in your own browser, one key per provider, and stays there between visits. Supported providers are Claude, OpenAI, DeepSeek and Kimi.
When you ask for an AI report, the key travels with that request so the call can be made on your behalf, then it is used and dropped. It is never written to a database and never logged, and there is no account for it to be attached to. Galley holds no key of its own.
Your provider bills you directly at their published rates. Galley adds nothing on top and takes nothing.
Not yet. Accounts and a saved history of your experiments are the next thing being built, so you will be able to keep a record of what you ran, what the verdict was, and what you decided.
Everything that works today will keep working without an account.
Yes. The report exports to PDF, which is usually the format that survives a forward to a client or a stakeholder who was not in the conversation. The verdict sentence sits at the top, so the person reading it does not need to interpret anything.
Nothing today. There is no account, no trial, and no card. If paid features arrive later, the analysis you can run now will stay available.
Bring the export and find out what it actually said.
Analyse a test