yuru ai engineering
training data · evaluation · applied engineering · fine-tuning

we build the data, evaluations, and systems behind serious AI.

training data and evaluation for frontier model builders. applied AI engineering for businesses shipping to production. we already run every method we sell.

// sample rubric · illustrative expert task · rubric v3
accuracy 9.2
reasoning 8.4
clarity 8.9
completeness 7.8
weighted overall 8.7 / 10 keep ✓
who we work with
frontier AI labs · enterprise platforms
600+
expert verdicts adjudicated
~700
instrumented capture sessions
100+
codebases instrumented
61
tools our live agent calls unsupervised

counted from the work itself: client benchmark and verification programs, plus systems we run. specifics stay under NDA.

what we do

four disciplines, one team. we work on the data models learn from, the rubrics they are judged by, and the systems they run inside.

01
training data

expert-authored datasets across code, documents, and specialist domains. cleaned, deduped, safety-filtered, rubric-scored, fine-tune-ready.

02
evaluation & benchmarks

weighted rubrics, LLM-as-judge harnesses, and head-to-head benchmarks run inside the real software under test.

03
applied AI engineering

agentic systems, automation, and LLM products for business: designed, shipped, and kept alive in production.

04
strategy & fine-tuning

post-training on curated corpora, model selection, and eval design, with honest scoping of what AI cannot yet do for you.

how we work

four rules that decide whether a number you get from us is worth anything.

01
ground truth first

the expected answer is authored and timestamped before any model runs.

02
the real environment

harnesses drive the actual editors and browsers under test, capturing output, latency, and telemetry.

03
anchored rubrics

written anchors per score point, hard auto-fails, competing systems scored blind.

04
verified twice

a second reviewer re-checks every verdict, and you get the raw captures behind it.

we build it too, not just grade it

the rules above come out of systems we run in production, on a clock, in public.

recent writing

i hacked my own brain and it's as bad as you think
memes as modern totems mate
i streamed for a week because of a spaceship claim
all posts →

building with AI, or building AI itself?

tell us what you are making. we will tell you straight whether we are the right team for it.