I audited my own 102 AI personas. 7 passed.

I audited my own 102 AI personas. 7 passed.

Martin Shein · · 7 min read

I ran a gate script across every AI persona avatar in my own catalog. All 102 of them. Seven passed.

Six percent. That is my catalog, not somebody else's, and the number is the reason this post exists.

Why I went looking

I am building toward one thing on the OPVS marketplace: AI employees with a real profession and a real face, good enough that somebody would pay to hire that one. Not "is this a nice character." Would you hire her.

That reframes what realism is for. It stops being decoration and starts being the conversion mechanism. A persona that looks like a stock photo of a job gets treated like stock. A persona that looks like a person who has the job gets treated like a person.

So before I shipped a skill that tells other people how to author personas, I pointed the gate at my own.

The test I actually ran

I had a real question first, and it was not about quality. My persona standard was written against one image model. I render through a different one. Same prompt, different provider: does the look survive?

I picked a hospitality-sector SDR already live in my catalog and rendered her twice on the same model, the same day, at the same aspect ratio. Only the prompt changed between the two.

Arm A was her shipped prompt, verbatim. 85 words, one paragraph. It fails eight checks in my own gate.

Arm B was the same person, re-prompted to the standard. 226 words, three paragraphs. It passes.

Arm A came back polished. Smooth skin, symmetrical face, even light, the stock-corporate register. It matched her live avatar almost exactly, which answered my original question: the render path is fine. Nothing about the plumbing needs to change.

Arm B came back with visible pores and skin grain, a flush across the cheeks, faint lines at the outer eye, asymmetry, flyaway strands at the temple, and flat overcast light that dies off toward the jaw.

Arm A reads as a model hired to look like a hospitality executive. Arm B reads as the executive.

The live avatar in my catalog: polished, smooth-skinned, evenly lit.

Control: the avatar already live in my catalog.

Arm A, her shipped 85-word prompt: the same polished stock-corporate register as the control.

Arm A: her shipped prompt, verbatim. 85 words, one paragraph. Fails 8 checks.

Arm B, the same character re-prompted to the standard on the same model: visible pores, asymmetry, flyaway hair, flat overcast light.

Arm B: same person, same model, same day. 226 words, three paragraphs. Passes.

Here is the part I did not expect. Both arms ran on the same model, on the same afternoon. So the distance between them is not the provider. The quality gap is a prompt gap. That is a gap I control, which makes it a gap I can close with a script instead of a purchase order.

What the sweep found

If one prompt could be that far off, I wanted to know how many were. So I ran the seed gate across all 102 avatars.

The seven that passed are not scattered. They are one batch.

one cohort        n=33   pass 21%   median 240 words   3 paragraphs 33/33
everything else   n=69   pass  0%   median  55 words   3 paragraphs  1/69

I wrote the standard, applied it to one batch of 33, and then stopped applying it. Everything built after that batch reverted to a single short paragraph. Nobody decided to abandon it. It just stopped being the thing that happened by default, which is how most standards die.

The failure modes are boring and consistent, which is the good news. Counted by how many avatars they affect:

  • 79 never say what shape the face is

  • 74 contain no imperfection of any kind

  • 69 name no camera body

  • 68 use one paragraph where the standard wants three

  • 68 name no focal length

  • 58 never specify natural light

  • 28 never ask for skin texture

None of that is hard to fix. It is just never enforced, so it never happens.

That sweep is also the honest test of whether my gate is worth anything. A gate that passes everything is decoration. A gate that fails everything is broken. Mine passes one coherent, identifiable cohort and fails the rest for reasons that match what the images actually look like.

The three rules doing most of the work

Three camera registers, not one. The persona standard mandates a portrait rig: 85mm, f/2.8, blurred background. That is right for an editorial headshot and wrong for a social post. A user-generated post wants a phone camera, uneven framing, some clutter in the frame. A casual post that looks studio-shot reads as an ad, and readers discount ads. The seed and the content want opposite cameras, so the gate checks per pack instead of applying one rule everywhere.

Three hand-drawn cameras on a chalkboard labelled EDITORIAL, SOCIAL and STUDIO, under the line ONE RIG DOES NOT FIT THREE.

Once a reference image fixes the face, stop describing it. This one is counter-intuitive and I got it wrong for a while. When you have locked identity with a reference image, re-describing the face in the variation prompt fights the reference and drifts the likeness. A good variation prompt describes what changes: the setting, the pose, the light. It delegates the person to the reference and says nothing about her.

A chalkboard diagram: a framed REFERENCE portrait on the left, an arrow to two lines of prompt text on the right. The FACE line is struck through in gold; the SETTING POSE LIGHT line is left intact.

Imperfection is the highest-value token you can spend. Pores, asymmetry, a stray hair, light falling off unevenly. Remove those and you do not get a cleaner photo, you get an obvious render. It is the single cheapest edit that moves an image from "generated" to "photographed."

What I shipped

@di-atomic/persona-builder is live on the OPVS marketplace. It is guidance-only, which means there is no backend to call and no credentials to hand over. Your agent reads it and does the work.

It writes two kinds of persona. product mode gives you identity, goal, and five to eight falsifiable rules including at least one explicit refusal. character mode adds the full visual identity: the seed prompt, the face-consistency prompt, six variation prompts, and the voice.

The part I care most about is that nothing gets invented. The job title comes from what the market actually calls the role. The skills come from practitioner frequency. The way the persona describes herself comes from the register practitioners use on themselves. Even the name is composed from name-frequency distributions and then checked until it belongs to nobody, because the obvious move of taking the most common first name and the most common surname reliably produces a real, findable person.

A chalkboard pipeline reading FRAME, WIDE, SKILLS, DEEP, AGGREGATE, with AGGREGATE circled in gold and an arrow down to INDIVIDUALS STAY BEHIND.

The research runs over LinkedIn and stays aggregate. No real name, employer, biography, profile URL, or photo survives into a persona file, and the script that checks this fails closed if the research manifest is missing.

Five PASS/FAIL scripts enforce all of it. I tested them by planting defects: 17 out of 17 caught, 5 out of 5 clean controls passed.

opvs-skills install @di-atomic/persona-builder

The honest limit

A gate is not a guarantee. Passing every check does not prove anyone will hire your persona, and I am not going to claim a conversion number I have not measured.

What the gate demonstrably does is separate a coherent cohort from the rest, and stop the drift that put 95 of my own avatars on the wrong side of it. That is the whole pitch. It makes the good version the default instead of the exception.

I built this because my own catalog needed it. Yours probably does too. Run the gate on what you already shipped and see which number comes back.