ai evaluation · benchmarks · production rescue
it demos beautifully and fails with real users.
the hard part is not building it. it is knowing whether it works, and whether last week's change helped or quietly made it worse. we build the evaluation that answers that, then fix what it finds.
looking for a website or ongoing development work instead? that is priced flatly and lives on the
websites page.
two kinds of engagement
same discipline, different deliverable.
track a
for teams shipping AI
an AI feature in production or close to it. good in the demo, uneven with real users, and nobody can say by how much.
- an eval suite for your use case, built from your real traffic, so a model upgrade becomes a number instead of an argument.
- failure diagnosis. we run your system against cases built to break it and hand back where it breaks, with transcripts, ranked by cost.
- the fix. retrieval that cites instead of inventing, prompts that hold under pressure, agents that recover instead of looping.
- regression protection. the suite re-runs every release, so you find out before your users do.
- model selection decided on your data and cost envelope, not a leaderboard.
- honest scoping, including telling you a problem is not ready for AI.
track b
for frontier labs
training or grading a model, and needing verdicts at a bar your own team would defend in review.
- expert training corpora, fine-tune-ready, with per-record scores attached.
- evaluation rubrics. written anchors per score point, plus auto-fails for what a naive score would rank highest.
- LLM-as-judge harnesses, validated against human adjudication rather than trusted out of the box.
- head-to-head benchmarking inside the real software under test, scored blind.
- human adjudication at volume, with an independent second reviewer on every verdict.
- RLHF and preference data, with the reasoning behind each ranking.
how it works, and what it costs
three steps. stop after any of them. the first is deliberately cheap, so you do not take our word for it.
step 01
$750
about two weeks
the diagnostic
an eval set from your real cases, a baseline score, and a ranked list of what is failing. yours either way, and it credits in full against the build.
step 02
from $5,000
fixed scope · fixed price
the build
we fix what the diagnostic found and build the full suite. price agreed in writing first, 50% up front. you get the harness.
step 03
from $2,000/mo
ongoing · cancel any month
keep measuring
the suite re-runs on every release and model upgrade, and we fix what it catches. scope agreed monthly. no contract.
an eval built once decays. models change under you, prompts drift, and March's failure modes are not September's. step 03 exists because that is the shape of the problem, not because we wanted a retainer.
why we are the ones to do this
most of our work is evaluation for the companies training frontier AI models. counted from the work, not estimated.
600+
expert verdicts adjudicated
~700
instrumented capture sessions
100+
codebases instrumented
one program benchmarked frontier AI coding assistants head to head inside the editors they ship in, across more than a hundred purpose-built task repositories.
another verified AI answers against cited evidence across medicine, law, finance, and science. every verdict is backed by verbatim quotes, then re-checked by a second reviewer. specifics stay under NDA.
how we run evals
a benchmark you can trust runs in the environment the model works in, then gets adjudicated twice.
01
ground truth first
experts author the expected answer before any model output is seen, and the timestamp is checked. scoring against a spec, not a vibe.
02
run in real software
custom harnesses drive the actual editors, browsers, and platforms under test, capturing every output, its latency, and the telemetry behind it.
03
score on anchored rubrics
every dimension scored against written anchors. competing systems never see each other's results, and judges justify before they score.
04
verify twice
a gated self-audit, then an independent reviewer re-checks every verdict against verbatim evidence. nothing ships on a single opinion.
how we grade data
raw examples in, a quality-ranked training set you can fine-tune on directly out.
01
ingest
raw examples into a fine-tune-ready schema: messages plus full tool-call signatures.
02
clean & safety-filter
dedupe, strip generic responses, and remove unsafe content before a dollar is spent on grading.
03
score on a weighted rubric
every record graded on weighted dimensions by a judge that reasons first, with auto-fails for the failure modes that matter.
04
rank & select
read the distribution, keep the top band, hand back a ranked set with the scores that got it there.
what you receive
every engagement hands back the artifact and the machinery that made it.
the eval suite
your cases, rubric, and baseline scores, re-runnable against any future model.
the rubric
written anchors per score point, auto-fail criteria, and a guide a new reviewer could pick up cold.
the harness
the code that ran the benchmark, so you can re-run it without us.
the evidence
raw captures, telemetry, and verbatim quotes behind every verdict, so any score traces back to what produced it.
the questions everyone asks
what if the diagnostic says my system is fine?
then we say so instead of inventing work. rare, but real. you keep the eval set either way, so you can prove it stays fine.
do we keep the eval suite and the harness?
yes, at every step including the $750 one. code, rubric, cases, scores. we have no interest in being the only ones who can measure your system.
what do you need from us to start?
access, and roughly twenty real examples of what it should handle. failure screenshots, angry tickets, and the cases your team argues about are ideal. logs work too.
how is this different from using an off-the-shelf LLM-as-judge?
a judge you have not validated against human verdicts is a random number generator with good manners. the work is tuning it until it agrees with experts, then telling you the agreement rate.
we do not have real traffic yet. is this too early?
no. the eval set gets authored from your spec instead of from logs. that is how labs do it pre-launch, and it is cheaper than finding out after you ship.
what does the $2,000 a month include?
re-running the suite on releases and model upgrades, reporting what moved, and fixing what it catches. scope agreed at the start of each month. larger work gets quoted on top rather than silently slowing down.
will you sign an NDA?
yes, before you share anything sensitive. most of our work is already under one, which is why the programs above are vague about who they were for.
who actually does the work?
the people you talk to are the people who build it. no account manager, no offshore handoff, no junior learning on your budget. engagements are scheduled rather than queued, so timing is worth asking about early.
find out what your AI is actually scoring.
$750, about two weeks, and you keep everything. if you do not need us, we say so before you spend it.