a multidisciplinary AI engineering firm
yuru is a small, senior team working across the AI stack: expert training data, benchmarks, and evaluation for teams building frontier models, applied engineering and integration for businesses, and autonomous systems of our own. our datasets, rubrics, and adjudicated verdicts support teams at frontier AI labs; the specifics stay under NDA. we are not a body shop: we design the rubric, build the harness, run the pipeline, and stand behind every score.
as far as NDA allows: recent programs include head-to-head benchmarking of AI coding assistants inside the editors they ship in, with custom instrumentation capturing every suggestion and its latency, and evidence-grounded verification of AI answers over expert documents in fields like medicine, law, and finance, every verdict backed by verbatim quotes and a second review.
we also build and run our own autonomous AI in production. that is where our methods come from, and why the work we sell is proven before it reaches you, not theorized on a slide.
what we do
training data
expert-authored and captured datasets across code, documents, and specialist domains. cleaned, deduped, safety-filtered, rubric-scored, and delivered fine-tune-ready.
evaluation & benchmarks
weighted rubrics, LLM-as-judge harnesses, and head-to-head benchmarks run inside the real software under test. ground truth authored first, anchored scoring scales, hard auto-fail criteria, a second reviewer on every verdict.
applied AI engineering
agentic systems, automation, and LLM products for business: designed, integrated, shipped, and kept alive in production. from harnesses that drive real desktop and web apps to browser agents and real-time voice.
strategy & fine-tuning
post-training on curated corpora, model selection, and eval design, with honest scoping of what AI can and can't yet do for your organization.
how we work
data: a pipeline, not a purchase
raw examples in; a quality-ranked training set out. we normalize to a fine-tune-ready schema, clean and safety-filter, score every record on a weighted rubric with an LLM-as-judge that reasons before it scores, then rank by the full score distribution and keep the top band, with per-record scores attached.
evals: run like experiments
ground truth authored before any model output is seen. custom harnesses drive the real editors, browsers, and platforms under test and capture every output with its latency. scoring is per-dimension against written anchors, competing systems scored independently, and a second reviewer re-checks every verdict against evidence quoted verbatim.
we build it too
burnt melba
a fully autonomous AI vtuber that live-streams unsupervised, drives real apps by browser automation, remembers, and sets its own goals. our proving ground for agent reliability, and where the data pipeline above was forged: 26 corpus versions generated, judged, and ranked so far.
yume
our native desktop app for agentic coding: parallel agents, deep shortcuts, multi-provider. a product we ship and support.
building with AI, or building AI itself?
tell us what you're making. we'll tell you straight whether we're the right team for it.
get in touch