134Sources 87,649Indexed items 12,625Full transcripts 55% → 8%Stub rate
Transcripts per source · 121 sources with captured text · log scale

Your AI demo works. Production is a different problem.

I take LLM and agent systems from prototype to something your team can run unattended — with evaluation, failure handling, and a handover so you own it after I'm gone.

Lund, Sweden · EU-based · working worldwide

Evidence

What's actually running.

One ingestion pipeline, unattended, against 134 third-party sources that each fail differently and mostly fail quietly. The numbers below come from my own systems and are reproducible from the repository — not from a case study I can't show you.

134
Data sources
RSS, podcast feeds, sitemaps, transcript APIs
87,649
Indexed items
Across every source, deduplicated
12,625
Full transcripts
910 MB of captured text
55% → 8%
Stub rate
Every remaining gap has a named, unfixable cause
18 h
Longest silent hang found
A TLS read urllib's timeout doesn't cover
15
Autonomous work units
Each with its own state file and decision log

Two bugs from that build, both invisible in a demo:

A TLS read that hung a worker for 18 hours — urllib's timeout covers connect, not read, so it needs a SIGALRM hard cap. And a paginated archive API that silently ignored offset and returned the same batch forever, which looks exactly like success until you count distinct IDs.

This is the work. Not the prompt.

ersiai OS Internal

The agent fleet that runs this studio. Scans deterministically and cheaply — no LLM in the hot loop — escalates only what needs judgment, logs every autonomous decision for audit.

OneShot Apply Live

An AI job-application studio: role discovery, tailored resumes, fit analysis, mock interviews. In public, with real users and real failure modes.

Dacheng Fayuan Live

A study platform for Mahayana Buddhism. Shipped and maintained for a non-English audience, which is its own class of problem.

FlowTranscript Beta

Transcription at length — the pipeline behind the corpus above, packaged as a product.

CoolVideoLab Beta

Short-form video, script to rendered clips as one reproducible pipeline rather than a pile of prompts.

CoolImageLab Beta

Image generation and editing with a reviewable workflow, so you can tell why an output changed.

Engagements

Choose the right scope.

Three ways to work together. Scope and price are agreed in writing before anything starts.

Audit

For teams with an LLM feature in progress — or stalled — who want to know if it's the right bet before spending another sprint.

€600
fixed price
  • Half-day session on your codebase and workflows
  • Written report: where agents pay off, and where they don't
  • Prioritized roadmap you can execute without me

If the honest answer is “don't build this”, that's what the report says.

Get started →

Sprint

For teams who know what they want but don't have someone who has shipped this kind of system before.

€3,000
starting at
  • 2–4 weeks, one scoped deliverable in production
  • Agent workflows, LLM pipelines or automation
  • Tests, docs, and a handover session
Get started →

Advisory

For teams who have shipped and want a second pair of eyes before the next architecture decision.

Retainer
monthly
  • Architecture reviews and unblocking sessions
  • Async access for your team's AI questions
  • Priority scheduling for follow-up sprints
Get started →
About

Evaluation is the job.

Most LLM products don't fail because the model is weak. They fail because nobody defined what “working” means, so nobody notices when it stops.

I spent a decade in machine learning and medical-imaging research — a field where a model that's 95% accurate can still be unusable, and where you have to show why it works, not just that it worked once. That's the discipline I bring to LLM products.

ersiai is a one-person studio, run alongside its own agent fleet — the same kind of system I build for clients.