The Product

Everything SkillMark measures.

One engine, two jobs: hire the right people, and grow the ones you already have. Here is the whole thing — both modes, every feature, and every role we assess.

One Engine, Two Modes

Hire externally. Develop internally.

The same engine, the same tasks, the same four axes — pointed either outward at people you are considering, or inward at the team you already have. No competitor productizes the second one.

Hiring mode

Send invite-token assessments to external candidates and compare them apples-to-apples in an isolated sandbox.

Internal L&D mode

Employees take assessments under their own identity. Track AI-proficiency progression over time with team analytics.

Manager-assigned growth

Managers assign tasks to team members and watch skills develop — internal AI-proficiency development nobody else offers.

The Problem

Nobody can see the skill that now decides the outcome.

Every professional uses AI daily — the ones you are interviewing and the ones already on your payroll. Hiring processes either ban it or can't read it, and for your existing team nobody is even trying to look.

01

AI is the new baseline

Every role uses AI now — from coding to writing to analysis. Banning AI in assessments tests the wrong skill.

02

You only ever see the output

A finished deliverable tells you nothing about how it was reached. That is true of a take-home from a candidate and equally true of the work your team shipped last sprint.

03

Nothing is comparable

Without a standard task and a standard rubric there is no fair way to compare two candidates, two teams, or the same person six months apart.

04

Domain knowledge is invisible

You can't tell whether someone knew which question to ask, or just accepted whatever the model said first — and that is the difference you are hiring and promoting on.

The Platform

See exactly how anyone works with AI.

01

Pick a role & task

Choose from 90+ ready-to-send assessments covering every role you hire for — engineering, product, marketing, design, QA, DevOps, data engineering, finance, legal, and more. Or author your own.

02

Candidate works with AI

Candidates and employees alike get a purpose-built workspace with an AI assistant, reference materials, and a timer. Everything is captured — prompts, decisions, output.

03

Review the full picture

Get the 4D fluency vector — Delegation, Description, Discernment, Diligence — plus per-dimension evidence for domain knowledge and output quality. Every rating cites the moment in the session that earned it.

Features

What no one else combines.

Not just what candidates produce — verified evidence of how they work with AI.

4D fluency vector

Four axes, not one number

Every session resolves to a fluency vector: Delegation (how much was handed over), Description (how clearly it was instructed), Discernment (whether the output was verified), Diligence (how efficiently). The vector is the result, not the single number derived from it.

Every keystroke

Full session replay

Screen recording with event timeline overlay. See every keystroke, every prompt, every decision — reviewable at any speed.

Multi-provider

Works with Claude & GPT

Multi-provider by design — Anthropic and OpenAI. Not a Claude-only point tool, so you see how candidates work with the AI they actually use.

Evidence-gated

Explainable, evidence-gated scoring

Penalty mode starts from 100 and subtracts named deductions — −12 vague prompts, −10 missing context — organized under the 4D fluency lens: Delegation, Description, Discernment, Diligence. Every deduction is gated on a real evidence quote from the session. Toggle alongside the weighted AI-Q score per assessment.

Every role

Coding & document tasks

Code editor for engineers. Rich text editor for PMs, marketers, designers. Each role gets the right workspace.

Real-world tasks

Reference materials & data

Provide CSVs, documents, briefs — candidates analyze real data with AI help, just like the actual job.

Why Trust Us

Depth that's shipped — not promised.

We're pre-launch and honest about it. Here's the verifiable substance already running in the product.

Multi-provider by design

Anthropic and OpenAI both supported today — not a Claude-only point tool.

A fluency vector, not a single score

Delegation, Description, Discernment, Diligence — reported as four axes so a strong average can't hide a weak one. A scalar exists only as a labelled projection for sorting.

Full synchronized replay

Code, AI chat, and screen recording on one timeline — every prompt and decision reviewable.

Evidence-gated, explainable scores

Every rating is backed by validated evidence quotes from the candidate's real session. Penalty mode surfaces itemized deductions per rubric criterion — fully explainable, not a black box.

Beyond Engineering

AI skills assessment for every role.

Not just coding. Candidates write PRDs, analyze data, draft emails, review designs, prepare budgets — all with AI assistance, all observable.

Engineering

19 tasks

Product Management

7 tasks

ML / Data Science

7 tasks

Product Design

6 tasks

Operations

6 tasks

Sales / BDR

6 tasks

Data Engineering

6 tasks

DevOps / SRE

5 tasks

Engineering Management

5 tasks

Finance / FP&A

5 tasks

Marketing

5 tasks

Recruiting / HR

5 tasks

Content / Writing

3 tasks

Analyst / BI

2 tasks

Customer Success

2 tasks

Developer Relations

1 tasks

Legal / Compliance

1 tasks

QA / Test Engineering

1 tasks

Solutions Architect

1 tasks

The Real Cost

The real cost isn't AI technology. It's people using AI inefficiently.

Tokens are cheap. Salaries are not. The cost of a team that delegates to AI without judgement never lands on an invoice — it lands in rework, in review time, and in work that quietly isn't as good as it looks. And it scales with you: every hire multiplies it, and nothing on your P&L tells you it happened.

Rework you never see coming

Output accepted from AI without checking looks finished. The cost lands later — in review, in QA, or in production — and by then nobody traces it back to the prompt that caused it.

You pay in salary, not tokens

An hour of an engineer's time costs orders of magnitude more than the inference they spend in it. Underused AI is a payroll problem wearing a tooling costume.

It grows faster than you do

Two people on the same salary with the same model produce very different work. At ten people that is an annoyance. At two hundred it is a budget line you never approved — and it compounds, because unmeasured habits get taught to every new joiner by the person sitting next to them.

You can't coach what you can't see

Reviewing a finished deliverable tells you nothing about how it was reached. Seeing the prompts, the pushback, and the discarded turns is what makes AI skill teachable instead of innate.

FAQ

Common questions

See how your next hire actually works.

Every prompt, every decision, every output — before you make an offer.

Invite-only access. Book a demo to get started.