My best CRO win was +25%. The retest said minus 14%.
Nothing broke. Both numbers were real. I just treated one run as a fact, wrote it into the client’s baseline, and every decision after that stood on it. cro-optimizer is the third skill in my CRO chain and the only one whose job is to disagree with you — it reads the outcome after a fix ships and decides whether the result survives two gates before anything gets banked.
For teams already running tests — who have noticed that the wins in the deck never quite show up in the revenue.

opvs-skills install @di-atomic/cro-optimizerThen say: “is this a real win?” or “the pricing test came back +18%”.
Why your CRO program keeps beating its forecast
test ends → +25% → "nice, ship it" baseline: +25% recorded deck: "+25% uplift" next test: built on that prior (no duration check, no power check, no guardrail metric, no discount — and no error)
Half of all tests are statistically null, so most weeks the honest answer was “nothing happened.” Recording it as +25% pollutes every audit that follows.
test ends → +25% → provenance: deploy v412, 9 days significance gate .... 5/5 PASS causal gate .......... 5/5 PASS guardrail: trial-to-paid flat discount applied VERDICT won banked: ~+12-18%, not +25%
The number that reaches your baseline is the one that survived both gates and the winner’s-curse haircut. It is smaller. It is also true.
A program that believes its own wins isn’t learning. It’s compounding its own noise.
What you actually get
Two gates, run on every result
The significance gate checks the sig floor, duration, power against a pre-declared MDE, peeking and thin samples. The causal gate checks the discount, certainty language, estimate stability, confounds and a named guardrail metric. Both must pass before a result is allowed to be a verdict.
The winner’s-curse haircut
Test wins overstate production by 20–50%, structurally — you picked that variant because its noise pointed up. The raw number is a ceiling, never an estimate. What gets written into your baseline is discounted to 50–80% of raw, every time.
Losers get reverted, not noted
Roughly 1 in 14 shipped variants actively hurts. On a lost verdict the next action is a rollback through SpiderPublish, not a line in a doc. And a lift that came from adding an absolute claim on a REACH, CLP or SDS page is flagged as a liability, never banked.
The two mechanics that do the work


One result, start to verdict
> "the pricing-page test came back +18%, ship it?"
precondition ........... analytics wired (PostHog) OK
content_deploy_status .. v412, shipped 9 days ago provenance fixed
conversions per variant pulled from PostHog
significance gate ...... 5/5 PASS (p<.05, 9d/2 cycles, 80% power, no peeking)
causal gate ............ 5/5 PASS (segments split, no novelty, guardrail named)
guardrail: revenue per visitor .... flat, not down
VERDICT won
discounted_lift ~+9-14% expected in production, NOT +18%
next_action keep deploy → write prior to cro-audit baseline
→ generateNextHypothesis: iterate the winner
> same call, a week earlier, day 3 of the test
significance gate ...... FAIL duration 3d < 7d, peeked at first significance
VERDICT inconclusive — no baseline write, keep running
The second block is the one that saves money. On day 3 the honest answer is “you do not know yet”, and the skill will not be talked out of it.
Where it stops

I built this by dogfooding the same stack I run for clients.
cro-optimizer is one skill in the system behind Di-Atomic — the marketing & compliance agency that runs cognitoAI, SpiderIQ and OPVS. If you want a conversion program whose reported wins you can actually spend against, operated by an agent that will tell you when a result is noise — in any of our seven languages, including regulated categories like REACH and CLP — that’s the day job. Let’s talk.
Book a 30-min call with Di-AtomicJust want the skill? Install @di-atomic/cro-optimizer free.