yuru web & ai engineering
about

you deal directly with whoever builds it.

yuru.be is an engineering studio in Tampa, Florida. the work splits two ways. we build and look after websites and software for local businesses and tech companies, flat priced, with the maintenance included rather than sold back to you. and we build training data and evaluation programs for frontier AI labs, the companies training the models everything else runs on. those specifics stay under NDA.

both come down to the same question: does it hold up when a real person uses it. we design the rubric, build the harness, run the pipeline, and hand over the evidence. nothing gets handed off after the quote.

recent programs

two of them, described as far as NDA allows.

one program benchmarked frontier AI coding assistants head to head inside the editors they actually ship in, capturing every suggestion with its latency and telemetry. expected outputs were authored and timestamped before any model ran, across more than a hundred task repositories, and competing tools never saw each other's results.

another verified AI answers against cited evidence in expert documents spanning medicine, law, finance, and science. every verdict is backed by verbatim quotes, checked against the exact page they came from. outside knowledge may never bridge a gap, and every record passes an independent second review.

we also run our own autonomous AI in production. that is where the methods come from.

600+
expert verdicts adjudicated
~700
instrumented capture sessions
100+
codebases instrumented
9
expert domains covered

how we think about the work

ground truth before output
the expected answer is written and timestamped before any model runs. scoring after seeing the answer is a rationalization.
nothing ships on one opinion
a gated self-audit, then an independent reviewer. we would rather find our own mistakes than have you find them.
read the distribution
a pass mark tells you nothing. the distribution tells you what to keep, what to cut, and what your rubric actually rewards.
test in the real environment
a model that works in a notebook and fails in the product was never measured properly. we drive the real software, which is slower to build than a notebook harness.
show the evidence
every score traces back to the capture or quote that produced it. you get the harness too.
say the unwelcome thing
including when a problem is not ready for AI, when a benchmark measures the wrong thing, or when you should not hire us.

what our own systems taught us

every method here has a version of itself running in one of these.

01

a desktop app for agentic coding that we ship, sell, and support. paying customers are an unforgiving check on claims about what agentic tooling can actually do.

02

a market system that grades its own predictions against reality on a schedule. it taught us to compute every number in code before the model sees it, and to publish the score instead of the claim.

03

an autonomous AI vtuber running unsupervised in public. she taught us what agent reliability costs when there is no operator to unstick anything, and our weighted-rubric grading pipeline was built to give her a voice worth listening to.

still deciding if this is a fit?

describe the problem and we will tell you whether it should be us, and what it would take.