I take LLM and agent systems from prototype to something your team can run unattended — with evaluation, failure handling, and a handover so you own it after I'm gone.
Lund, Sweden · EU-based · working worldwide
One ingestion pipeline, unattended, against 134 third-party sources that each fail differently and mostly fail quietly. The numbers below come from my own systems and are reproducible from the repository — not from a case study I can't show you.
Two bugs from that build, both invisible in a demo:
A TLS read that hung a worker for 18 hours — urllib's
timeout covers connect, not read, so it needs a SIGALRM hard
cap. And a paginated archive API that silently ignored offset
and returned the same batch forever, which looks exactly like success until
you count distinct IDs.
This is the work. Not the prompt.
The agent fleet that runs this studio. Scans deterministically and cheaply — no LLM in the hot loop — escalates only what needs judgment, logs every autonomous decision for audit.
An AI job-application studio: role discovery, tailored resumes, fit analysis, mock interviews. In public, with real users and real failure modes.
A study platform for Mahayana Buddhism. Shipped and maintained for a non-English audience, which is its own class of problem.
Transcription at length — the pipeline behind the corpus above, packaged as a product.
Short-form video, script to rendered clips as one reproducible pipeline rather than a pile of prompts.
Image generation and editing with a reviewable workflow, so you can tell why an output changed.
Three ways to work together. Scope and price are agreed in writing before anything starts.
For teams with an LLM feature in progress — or stalled — who want to know if it's the right bet before spending another sprint.
If the honest answer is “don't build this”, that's what the report says.
Get started →For teams who know what they want but don't have someone who has shipped this kind of system before.
For teams who have shipped and want a second pair of eyes before the next architecture decision.
Most LLM products don't fail because the model is weak. They fail because nobody defined what “working” means, so nobody notices when it stops.
I spent a decade in machine learning and medical-imaging research — a field where a model that's 95% accurate can still be unusable, and where you have to show why it works, not just that it worked once. That's the discipline I bring to LLM products.
ersiai is a one-person studio, run alongside its own agent fleet — the same kind of system I build for clients.