The Product
One engine, two jobs: hire the right people, and grow the ones you already have. Here is the whole thing — both modes, every feature, and every role we assess.
One Engine, Two Modes
The same engine, the same tasks, the same four axes — pointed either outward at people you are considering, or inward at the team you already have. No competitor productizes the second one.
Send invite-token assessments to external candidates and compare them apples-to-apples in an isolated sandbox.
Employees take assessments under their own identity. Track AI-proficiency progression over time with team analytics.
Managers assign tasks to team members and watch skills develop — internal AI-proficiency development nobody else offers.
The Problem
Every professional uses AI daily — the ones you are interviewing and the ones already on your payroll. Hiring processes either ban it or can't read it, and for your existing team nobody is even trying to look.
Every role uses AI now — from coding to writing to analysis. Banning AI in assessments tests the wrong skill.
A finished deliverable tells you nothing about how it was reached. That is true of a take-home from a candidate and equally true of the work your team shipped last sprint.
Without a standard task and a standard rubric there is no fair way to compare two candidates, two teams, or the same person six months apart.
You can't tell whether someone knew which question to ask, or just accepted whatever the model said first — and that is the difference you are hiring and promoting on.
The Platform
Choose from 90+ ready-to-send assessments covering every role you hire for — engineering, product, marketing, design, QA, DevOps, data engineering, finance, legal, and more. Or author your own.
Candidates and employees alike get a purpose-built workspace with an AI assistant, reference materials, and a timer. Everything is captured — prompts, decisions, output.
Get the 4D fluency vector — Delegation, Description, Discernment, Diligence — plus per-dimension evidence for domain knowledge and output quality. Every rating cites the moment in the session that earned it.
Features
Not just what candidates produce — verified evidence of how they work with AI.
Every session resolves to a fluency vector: Delegation (how much was handed over), Description (how clearly it was instructed), Discernment (whether the output was verified), Diligence (how efficiently). The vector is the result, not the single number derived from it.
Screen recording with event timeline overlay. See every keystroke, every prompt, every decision — reviewable at any speed.
Multi-provider by design — Anthropic and OpenAI. Not a Claude-only point tool, so you see how candidates work with the AI they actually use.
Penalty mode starts from 100 and subtracts named deductions — −12 vague prompts, −10 missing context — organized under the 4D fluency lens: Delegation, Description, Discernment, Diligence. Every deduction is gated on a real evidence quote from the session. Toggle alongside the weighted AI-Q score per assessment.
Code editor for engineers. Rich text editor for PMs, marketers, designers. Each role gets the right workspace.
Provide CSVs, documents, briefs — candidates analyze real data with AI help, just like the actual job.
Why Trust Us
We're pre-launch and honest about it. Here's the verifiable substance already running in the product.
Anthropic and OpenAI both supported today — not a Claude-only point tool.
Delegation, Description, Discernment, Diligence — reported as four axes so a strong average can't hide a weak one. A scalar exists only as a labelled projection for sorting.
Code, AI chat, and screen recording on one timeline — every prompt and decision reviewable.
Every rating is backed by validated evidence quotes from the candidate's real session. Penalty mode surfaces itemized deductions per rubric criterion — fully explainable, not a black box.
Beyond Engineering
Not just coding. Candidates write PRDs, analyze data, draft emails, review designs, prepare budgets — all with AI assistance, all observable.
Engineering
19 tasks
Product Management
7 tasks
ML / Data Science
7 tasks
Product Design
6 tasks
Operations
6 tasks
Sales / BDR
6 tasks
Data Engineering
6 tasks
DevOps / SRE
5 tasks
Engineering Management
5 tasks
Finance / FP&A
5 tasks
Marketing
5 tasks
Recruiting / HR
5 tasks
Content / Writing
3 tasks
Analyst / BI
2 tasks
Customer Success
2 tasks
Developer Relations
1 tasks
Legal / Compliance
1 tasks
QA / Test Engineering
1 tasks
Solutions Architect
1 tasks
The Real Cost
Tokens are cheap. Salaries are not. The cost of a team that delegates to AI without judgement never lands on an invoice — it lands in rework, in review time, and in work that quietly isn't as good as it looks. And it scales with you: every hire multiplies it, and nothing on your P&L tells you it happened.
Output accepted from AI without checking looks finished. The cost lands later — in review, in QA, or in production — and by then nobody traces it back to the prompt that caused it.
An hour of an engineer's time costs orders of magnitude more than the inference they spend in it. Underused AI is a payroll problem wearing a tooling costume.
Two people on the same salary with the same model produce very different work. At ten people that is an annoyance. At two hundred it is a budget line you never approved — and it compounds, because unmeasured habits get taught to every new joiner by the person sitting next to them.
Reviewing a finished deliverable tells you nothing about how it was reached. Seeing the prompts, the pushback, and the discarded turns is what makes AI skill teachable instead of innate.
FAQ
Every prompt, every decision, every output — before you make an offer.
Invite-only access. Book a demo to get started.