Everyone takes the exact same test
One environment. One AI model, locked. Every candidate runs the identical setup, so you compare engineers, not who pays for the fancier AI subscription. Same starting line, every time.
The whiteboard is dead.
Every engineer now wields an AI, so output no longer tells them apart. We measure the one thing that still does — judgment.
Why we exist
Whiteboard puzzles and LeetCode never measured the job. And now that every engineer has an AI assistant, they measure it even less. The question isn't "can you write a binary search from memory." It's "can you direct an AI through a real, messy codebase and know when it's wrong?"
That's a skill. It's the skill. And until now, nobody tested for it.
How it works
What you get
One environment. One AI model, locked. Every candidate runs the identical setup, so you compare engineers, not who pays for the fancier AI subscription. Same starting line, every time.
Close the tab. Lose your wifi. Refresh by accident. It does not matter. Your progress saves continuously and the session resumes exactly where you left off.
A precise, line-by-line record of exactly what you built, the full diff of your work, right alongside your reasoning.
Perfect code is not the point. We watch how you scope the problem, steer Claude Code, verify its work, and recover when it goes wrong.
The task hides subtle pitfalls inside a realistic project, the kind that trip up anyone who trusts the AI blindly. Catching them is exactly what we look for.
One focused block with a legible timer that informs without inducing panic — and you always know exactly what is recorded, stated up front before the clock starts. Spacious, not stressful; no hidden surveillance.
Inside the assessment
The read is more than a transcript. It is a complete assessment — a verdict, the evidence behind it, and the moments that decided it, generated from the raw JSONL the moment an attempt closes.
Every pass and fail is decided in code against a fixed rubric; a model writes only the prose. So reviewers read judgment, not vibes.
Written on submit. No reviewer waits, no transcript gets read cold.
The verdict and every signal are computed, not improvised by a model.

Fairness and trust
Because every candidate gets the byte-for-byte identical environment and the same locked AI model, scores mean something. No environment drift, no "well, they had a better tool." Just a clean, repeatable signal.
Every assessment produces a complete, reviewable record: the reasoning, the decisions, the exact changes. Hiring calls backed by what someone actually did, not a gut feeling from 45 minutes on a call.
Reviewers grade against a clear rubric, with the option to hide candidate identity while scoring. Same bar for everyone.
The same attempt that gives reviewers a clean, comparable signal gives candidates an interview that finally looks like the actual job. One run, no trade-off.
Every candidate runs the byte-for-byte identical setup with one locked model, so scores compare engineers — not subscriptions. Grade an attempt in minutes from a clean evidence trail. Fair, repeatable, defensible.
Finally, an interview that looks like the work: a real codebase and real Claude Code, driven the way you actually drive it. Nothing to install, and nothing lost if the tab closes — judged on judgment, not memorized algorithms.
The same AI for everyone. You can't buy a better score.
Questions
The coding interview, rebuilt for engineers who work with AI.