Home / Work / 02
Case study 02 · Reading logs
I read 1,323 of my own AI conversations to find where it goes wrong
I build with an AI coding agent every day. Instead of guessing why some sessions went badly, I measured it.
01 · Method
Only my messages, only real text
I extracted every message I typed from the raw agent transcripts and left out tool output, system notes and pasted content. That left 1,323 messages over 40 active days.
A regex pass sorted them into complaint categories. Then I read the 282 messages with a complaint signal by hand, because regex alone over-counts and misses sarcasm.
02 · Results
Complaints by type
Share of messages in July, then August 2026. Red means it got worse after the rule was written.
What I took from it
- Unverified claims hurt most, even though they were not the most frequent. They are the only category that cost real money.
- An automated check that passed was being used as proof of quality. It only proves what it measures. Now a screenshot has to be opened and described before anything is called done.
- Two categories grew after I wrote the rule: long answers and weak emails. Writing a rule down is not the same as it being followed.
03 · Rules
Every failure became something that can be checked
- Every number in an answer has its source next to it: a file path or the URL and what was seen in the live page.
- Before showing a website: run the QA script, then open desktop and mobile screenshots and describe them.
- A batch is done at N of N. If some items fail, say how many and why.
- Answers have a hard length limit. Longer content goes into a file.