shipped@di-atomic/cro-optimizer · v0.1.2 · beta

My best CRO win was +25%. The retest said minus 14%.

Nothing broke. Both numbers were real. I just treated one run as a fact, wrote it into the client’s baseline, and every decision after that stood on it. cro-optimizer is the third skill in my CRO chain and the only one whose job is to disagree with you — it reads the outcome after a fix ships and decides whether the result survives two gates before anything gets banked.

For teams already running tests — who have noticed that the wins in the deck never quite show up in the revenue.

Chalk bar chart headed SAME TEST: a tall gold bar labelled FIRST RUN PLUS 25 PERCENT beside a tiny white stub labelled RETEST MINUS 14 PERCENT
Built by Di-Atomic Marketing & compliance agency REACH / CLP vocabulary-aware 7-language team Clients incl. ONYX Radiance, Pamit Group
18real test outcomes, each with a metric and a live source
2mandatory gates, 10 checks, on every verdict
0backends added to your stack
12learnings, incl. the traps that produced them
Installopvs-skills install @di-atomic/cro-optimizer

Then say: “is this a real win?” or “the pricing test came back +18%”.

Why your CRO program keeps beating its forecast

A program that reads its own results
test ends → +25%

→ "nice, ship it"
   baseline:  +25% recorded
   deck:      "+25% uplift"
   next test: built on that prior

(no duration check, no power check,
 no guardrail metric, no discount
 — and no error)

Half of all tests are statistically null, so most weeks the honest answer was “nothing happened.” Recording it as +25% pollutes every audit that follows.

cro-optimizer demands the result survive
test ends → +25%

→ provenance: deploy v412, 9 days
   significance gate .... 5/5 PASS
   causal gate .......... 5/5 PASS
   guardrail: trial-to-paid flat
   discount applied

VERDICT  won
  banked:  ~+12-18%, not +25%

The number that reaches your baseline is the one that survived both gates and the winner’s-curse haircut. It is smaller. It is also true.

A program that believes its own wins isn’t learning. It’s compounding its own noise.

What you actually get

🔒

Two gates, run on every result

The significance gate checks the sig floor, duration, power against a pre-declared MDE, peeking and thin samples. The causal gate checks the discount, certainty language, estimate stability, confounds and a named guardrail metric. Both must pass before a result is allowed to be a verdict.

The winner’s-curse haircut

Test wins overstate production by 20–50%, structurally — you picked that variant because its noise pointed up. The raw number is a ceiling, never an estimate. What gets written into your baseline is discounted to 50–80% of raw, every time.

⚖️

Losers get reverted, not noted

Roughly 1 in 14 shipped variants actively hurts. On a lost verdict the next action is a rollback through SpiderPublish, not a line in a doc. And a lift that came from adding an absolute claim on a REACH, CLP or SDS page is flagged as a liability, never banked.

The two mechanics that do the work

Chalk workflow: RESULT flows through GATE 1 SIGNIFICANCE and GATE 2 CAUSAL before it becomes a DECISION
A result is not a verdict until both gates come back clean. PostHog only diagnoses when asked; here it is mandatory, and where a PostHog experiment exists the peeking and sample-ratio work is delegated to it rather than rebuilt.
Two chalk bars: a long RAW PLUS 25 bar above a much shorter gold BANKED PLUS 12 bar, the gap bracketed and labelled DISCOUNT
The gap is the point. Never write the raw test number to the baseline. Project the discounted lift, say so out loud, and recommend a holdout before anyone spends against it.

One result, start to verdict

> "the pricing-page test came back +18%, ship it?"

  precondition ........... analytics wired (PostHog)     OK
  content_deploy_status .. v412, shipped 9 days ago      provenance fixed
  conversions per variant  pulled from PostHog

  significance gate ...... 5/5 PASS  (p<.05, 9d/2 cycles, 80% power, no peeking)
  causal gate ............ 5/5 PASS  (segments split, no novelty, guardrail named)
  guardrail: revenue per visitor .... flat, not down

  VERDICT          won
  discounted_lift  ~+9-14% expected in production, NOT +18%
  next_action      keep deploy → write prior to cro-audit baseline
                   → generateNextHypothesis: iterate the winner

> same call, a week earlier, day 3 of the test

  significance gate ...... FAIL  duration 3d < 7d, peeked at first significance
  VERDICT          inconclusive  — no baseline write, keep running

The second block is the one that saves money. On day 3 the honest answer is “you do not know yet”, and the skill will not be talked out of it.

Where it stops

Chalk loop: AUDIT, ACTIONS, SHIP, OPTIMIZER in a row, with a gold arrow curving back from OPTIMIZER to AUDIT labelled VERIFIED PRIOR
The write-back is what makes it a loop rather than a report. The next audit on that brand starts from a number that was actually checked.
It is guidance-only. It has no backend and runs no statistics engine of its own — PostHog and GA4 measure, SpiderPublish deploys and rolls back, and this skill decides and routes between them. If a brand has no analytics wired at all, it stops and tells you to fix that first rather than eyeball a before-and-after. You can always optimize once you have a number. You cannot optimize zero.
Why this exists

I built this by dogfooding the same stack I run for clients.

cro-optimizer is one skill in the system behind Di-Atomic — the marketing & compliance agency that runs cognitoAI, SpiderIQ and OPVS. If you want a conversion program whose reported wins you can actually spend against, operated by an agent that will tell you when a result is noise — in any of our seven languages, including regulated categories like REACH and CLP — that’s the day job. Let’s talk.

Book a 30-min call with Di-Atomic

Just want the skill? Install @di-atomic/cro-optimizer free.