Open to junior AI engineering roles · Brno, remote
I build AI agents and the tests that keep them honest.
Most of my time goes into the parts around the prompt: which tools the model gets, what it is not allowed to do, and reading logs until I find where it went wrong.
Simplified and translated from Czech. The agent has no tool that can change a price.
The number that changed how I build
The model sounded sure. It got four digits wrong.
A caller dictated a phone number in Czech words. Llama 3.3 70B wrote it down with full confidence and the wrong digits. Nothing in the reply hinted at a problem.
Since then: the model talks, code handles facts that must be exact.
Parser: 5/5 correct. Its result overrides whatever the model puts in the booking.
Selected work
Three projects, each with the failures left in
Every case study follows the same order: the problem, what I built, how I tested it, and what is still broken.
Czech phone agent for tradespeople
Takes job requests by phone, quotes from the real price list and books a slot. Safety by design, not by prompt.
Reading 1,323 AI conversations to find where it breaks
Parsed 116 raw agent transcripts, sorted every message and turned each failure type into a checkable rule.
Outreach pipeline with a linter on AI-written email
Lead scraping, a measured fact per company, LLM drafts and a linter built from every rejected draft.
Tools
Small tools, each built because something kept going wrong
A linter for AI email, website QA across screen widths, structured classification with probabilities, a model router with fallback, semantic search over ~5,600 notes.
What I can't do yet
- Most of my code is written by Claude. The ideas, decisions, tests and fixes are mine. Writing it unaided is what I'm working on.
- TypeScript: basics only. Next on my list.
- None of my agents has run in production with real users.