Goodfit CopilotGoodfit Copilot
Take Sample Assessment

The whiteboard is dead.

Coding interviews are broken.
We rebuilt it.

Every engineer now wields an AI, so output no longer tells them apart. We measure the one thing that still does — judgment.

Why we exist

The old interview died the day AI arrived.

Whiteboard puzzles and LeetCode never measured the job. And now that every engineer has an AI assistant, they measure it even less. The question isn't "can you write a binary search from memory." It's "can you direct an AI through a real, messy codebase and know when it's wrong?"

That's a skill. It's the skill. And until now, nobody tested for it.

How it works

From a real repo to a reasoning trail, in three steps.

01
Open a real codebase.
Not a toy puzzle. An actual brownfield project with real complexity, in a full browser IDE. Nothing to install.
02
Drive Claude Code.
Claude Code is ready and waiting. Give it the task. Direct it, question it, course-correct it, however you would really work.
03
We capture the whole story.
Every prompt, every decision, every change. Not just the final code, but the reasoning trail that shows how you got there.

What you get

The nice little things that add up to real signal.

Everyone takes the exact same test

One environment. One AI model, locked. Every candidate runs the identical setup, so you compare engineers, not who pays for the fancier AI subscription. Same starting line, every time.

Your work is never lost

Close the tab. Lose your wifi. Refresh by accident. It does not matter. Your progress saves continuously and the session resumes exactly where you left off.

Every change, tracked

A precise, line-by-line record of exactly what you built, the full diff of your work, right alongside your reasoning.

We measure judgment, not output

Perfect code is not the point. We watch how you scope the problem, steer Claude Code, verify its work, and recover when it goes wrong.

A codebase with real teeth

The task hides subtle pitfalls inside a realistic project, the kind that trip up anyone who trusts the AI blindly. Catching them is exactly what we look for.

Calm, and transparent by design

One focused block with a legible timer that informs without inducing panic — and you always know exactly what is recorded, stated up front before the clock starts. Spacious, not stressful; no hidden surveillance.

Inside the assessment

See exactly how someone works with an agent.

Steer a live agent.
Claude Code is ready in a real browser IDE. Give it the task, question its plan, course-correct it. We watch how you drive, not just where you land.
Graded on judgment, not output.
Modern agents write plausible code unprompted. We score how you scope the problem, verify the work, and catch the pitfalls a blind truster ships.
The read writes itself.
Every prompt and change becomes an auto-generated assessment: a verdict, the evidence, and the moments that decided it. Reviewed in code, not vibes.

The read brings the whole attempt together.

The read is more than a transcript. It is a complete assessment — a verdict, the evidence behind it, and the moments that decided it, generated from the raw JSONL the moment an attempt closes.

Every pass and fail is decided in code against a fixed rubric; a model writes only the prose. So reviewers read judgment, not vibes.

Auto-generated

Written on submit. No reviewer waits, no transcript gets read cold.

Decided in code

The verdict and every signal are computed, not improvised by a model.

The read — an auto-generated assessment over a candidate's transcript

Fairness and trust

For the people doing the hiring, fairness is the product.

Comparable by construction.

Because every candidate gets the byte-for-byte identical environment and the same locked AI model, scores mean something. No environment drift, no "well, they had a better tool." Just a clean, repeatable signal.

Evidence, not vibes.

Every assessment produces a complete, reviewable record: the reasoning, the decisions, the exact changes. Hiring calls backed by what someone actually did, not a gut feeling from 45 minutes on a call.

Built for low bias.

Reviewers grade against a clear rubric, with the option to hide candidate identity while scoring. Same bar for everyone.

Two sides, one signal

Built for the people hiring, and the people being hired.

The same attempt that gives reviewers a clean, comparable signal gives candidates an interview that finally looks like the actual job. One run, no trade-off.

  • Scoping
  • Steering
  • Verification
  • Recovery
  • Verdict

For hiring teams

Every candidate runs the byte-for-byte identical setup with one locked model, so scores compare engineers — not subscriptions. Grade an attempt in minutes from a clean evidence trail. Fair, repeatable, defensible.

For candidates

Finally, an interview that looks like the work: a real codebase and real Claude Code, driven the way you actually drive it. Nothing to install, and nothing lost if the tab closes — judged on judgment, not memorized algorithms.

Opus 4.8

The same AI for everyone. You can't buy a better score.

Questions

Anticipating the obvious objections.

Anyone can prompt. We measure who can steer.

The coding interview, rebuilt for engineers who work with AI.