Isolated sandbox. Verified evidence.

AI changed how
people work. Now measure it.

Candidates do real tasks with Claude or GPT in an isolated sandbox — never your real repo. You get verified evidence: every prompt, scored automatically on four axes — how much they delegate, how well they direct, whether they verify, and how efficiently they work. One engine for hiring and for growing your own team.

How It Works
Have an invite code?

Invite-only access. Book a demo and we'll get your team set up.

How It Works

One assessment replaces your entire evaluation process.

Your current process

  • Phone screen / initial call
  • Take-home assignment
  • Review meeting with team
  • Second round interviews
  • Team debrief & decision

2-4 timeline

Timeline

8+

Team time

>$500

Per candidate

With SkillMark

Recommended
  • Candidate completes a role-specific task with AI
  • Every prompt, edit, and decision captured
  • Evidence-backed scores on the 4D fluency vector — weighted or penalty mode
  • Hidden gems test domain knowledge
  • Full session replay available

0

Team time

30-60 min

Per candidate

Any

Role you hire for

4

fluency axes

Delegation, Description, Discernment, Diligence — a vector, not a single number, so a strong score can't hide a weak axis.

90+

ready-to-send assessments

Built and ready to send across every role you hire for — engineering, PM, marketing, design, QA, DevOps, data engineering, sales, CS, finance, legal, and more. Send one today, or author your own.

100%

captured

Every prompt, every edit, every decision — fully observable and replayable.

The Scoring Model

Four axes, not one number.

A single blended score lets a strong axis hide a weak one — and the two people it conflates need completely different answers from you.

Fluency vector

Illustrative session

Delegation78
Description84
Discernment71
Diligence66

Delegation is most readable next to Discernment: heavy delegation with strong discernment is expert work; heavy delegation without it is not.

Features

What no one else combines.

Not just what candidates produce — verified evidence of how they work with AI.

4D fluency vector

Four axes, not one number

Every session resolves to a fluency vector: Delegation (how much was handed over), Description (how clearly it was instructed), Discernment (whether the output was verified), Diligence (how efficiently). The vector is the result, not the single number derived from it.

Every keystroke

Full session replay

Screen recording with event timeline overlay. See every keystroke, every prompt, every decision — reviewable at any speed.

One Engine, Two Modes

Hire externally. Develop internally.

The same engine, the same tasks, the same four axes — pointed either outward at people you are considering, or inward at the team you already have. No competitor productizes the second one.

FAQ

Common questions

Still have questions? See all questions →

See how your next hire actually works.

Every prompt, every decision, every output — before you make an offer.

Invite-only access. Book a demo to get started.

The Films

The argument, in eight pieces.

Short films on what AI-era hiring actually misses — and one long essay on the eight-hundred-and-eighty seconds where it all goes wrong. Swipe through.

FilmTwo SessionsTwo candidates, identical output. The transcripts are where they stop looking alike.
FilmEveryone Looks CompetentWhen the AI writes the answer, the résumé stops telling you anything.
FilmNobody Checks the ChiefOne planted fact on page 22. Four executive seats, four ways of missing it.
FilmThe Unbilled LineThe P&L has every line except the one a bad hire actually costs you.
FilmMeasure the WorkWhat a session recording shows that a score never will.
FilmYes, Use the AIFor the candidate about to sit down: the tools are not the test.
FilmWhat's PlantedEvery brief carries something deliberately wrong. Catching it is the assessment.
EssayThe 880-Second WallA long-form post-mortem of the fourteen minutes that decide a hire.
Two Sessions1 / 8