Photo by MART PRODUCTION on Pexels
Headless commerce solved a set of problems that the themed storefront could not. You can build the front end the brand needs, ship it from your own repository, and keep Shopify behind the API for catalog, cart and checkout. What it did not solve is experimentation. The testing feature Shopify ships assumes an online store theme, and the moment the storefront moves to Hydrogen, Next.js or a custom stack, that assumption stops holding.
The limit is stated plainly on Shopify's requirements page for rollouts. Theme changes, and checkout and accounts configuration changes, apply only to your online store, and headless and custom storefront checkouts are not supported for those changes and cannot be tested with rollouts (Shopify Help Center).
Read that alongside the rollout types and the shape of the problem appears. A rollout in Shopify can change your online store theme, your checkout and accounts configuration, or your product catalogs, and the Experiment type, which compares a control against a treatment, needs the Grow plan or higher (Shopify Help Center). Every one of those levers is a property of the hosted storefront. On a headless build there is no theme to swap and no storefront checkout to reconfigure, so the native path runs out exactly where the interesting questions begin.
The reflex is to drop a client-side snippet into the page and call it done. On a headless storefront that approach fights the architecture:
The markup comes from your application, not from Shopify's renderer, so a DOM-level editor has to reconstruct what your components already control.
A CDN sits in front of the response and caches it, so the same cached HTML is served to visitors who are supposed to see different variants.
Assignment often lands after the page has painted, which shows one version and then swaps it. That flash is visible, it hurts Core Web Vitals, and it contaminates the result because the visitor's first impression was not the variant you are testing.
Price is written by your app from the Storefront API. If the experiment changes a price on the page and not in the cart, the visitor sees one number and pays another.
Before the numbers mean anything, five things have to be true.
Assignment happens before the first pixel renders. That usually means the server or the edge decides which variant to return, not a script that runs afterwards.
Assignment is sticky. A visitor who saw variant B on Monday has to see variant B on Thursday, which needs a stable identifier rather than a session cookie that expires with the tab.
The variant follows the visitor through the funnel. Cart, checkout, order and analytics all have to agree on which version produced the revenue.
One variable changes per experiment, with a metric and a sample size fixed before launch.
The cache is variant-aware. If the edge serves one response to everyone, a well-built test will read as a flat no-difference result rather than an error, which is the worst possible failure mode.
Client-side is the fastest to set up and the weakest to trust. It works for cosmetic changes such as copy, badges or image order, and it is a poor fit for anything that touches price, availability or checkout.
Server-side covers everything the application renders, including price and product data, and it is the only layer where a variant can be applied before paint. It costs engineering time, since the experiment becomes a branch in your own codebase.
Edge middleware sits between the two. The decision is made close to the visitor and the response is assembled with the assignment already applied, which removes the flash and keeps the origin out of it. The trade-off is that edge logic is a second place your variant definitions live, and it has to stay in step with the application.
A headless build buys you control over the front end, not a different answer to how often a test wins. In a meta-analysis of 1,001 A/B tests, 33.5% produced a statistically significant positive result, with a mean lift of 15.9% among the winners and a median of 7.5% (Analytics-Toolkit). Plan for a series of experiments rather than a decisive single one, and treat a test that reached no conclusion as a closed question.
Two practices matter more in a headless build than in a hosted store, because both are easier to get wrong there:
Guard against sample ratio mismatch. If your assignment logic has a bug, the split between variants drifts away from the configured percentage and the result is quietly wrong. Check the ratio before you read the outcome.
Do not peek. An edge or server-side test is often monitored from a dashboard that updates in real time, and reading it every morning until the number looks good turns a random fluctuation into a shipped decision.
Some teams should build this. If you already run feature flags, you have an assignment service, and your engineers can wire the Storefront API and the analytics pipeline into it, the experiment layer is a small addition to infrastructure you own.
The reason to buy instead is scope. Assignment, sticky bucketing, edge delivery, variant-aware caching, funnel-consistent pricing and a significance calculation with a rule for calling a winner is a product, not a sprint. Elevate's A/B testing for Shopify Headless Store connects to a storefront built on Hydrogen, Next.js or a custom stack, and sets experiments at the layers a headless build actually has: targeting by device, location or visitor behaviour, tests scoped to a campaign or a traffic source, and control over how much of the traffic enters each experiment. Reports land in real time with conversion rate, average order value and revenue per visitor per variant, and significance calculated automatically.
One detail to check before you assume it fits your budget: on the vendor's pricing page, headless support is listed as a Premium feature rather than something included from the entry plan, and price testing starts on the middle tier. Neither is unusual for this category, but the per-plan split decides the real monthly cost, so read the table rather than the summary line (Elevate pricing page, read 8 October 2026).
The list is short, and the order is decided by traffic volume.
On a store with a small audience, start with the product page: hero image, title, the position of add-to-cart, and the framing of delivery expectations.
On a store with real volume, price is the highest-leverage variable and the one most headless builds never test, because it has to be consistent end to end.
On a store running paid campaigns, scope variants to the landing experience for that campaign, where the traffic is densest and the intent is known.
Can Shopify's own experiments run on a headless storefront?
Not for theme or checkout configuration changes. Shopify's requirements page states that those changes apply only to the online store, and that headless and custom storefront checkouts are not supported and cannot be tested with rollouts.
Where should the experiment logic live?
Wherever the variant can be decided before the first render. In practice that means the application server or the edge, with the client kept for cosmetic changes only.
How much traffic does a headless A/B test need?
It depends on the baseline conversion rate and the size of the lift you want to detect. The rule that matters is setting the sample size before launch; a variant that leads after a few hundred visitors is a leading variant, not a winner.
Does testing slow the storefront down?
A client-side snippet that assigns after paint can add layout shift. Deciding the variant at the edge or the server removes that cost, which is one of the reasons server-side testing is the default choice on headless builds.
Discover our other works at the following sites:
© 2026 Danetsoft. Powered by HTMLy