Home / Work / 02

Case study 02 · Reading logs

I read 1,323 of my own AI conversations to find where it goes wrong

I build with an AI coding agent every day. Instead of guessing why some sessions went badly, I measured it.

When
August 2026
Data
116 raw transcripts, June to August 2026
Method
JSONL parsing, regex categories, manual reading
Output
Checkable rules, then tracking if they held

01 · Method

Only my messages, only real text

I extracted every message I typed from the raw agent transcripts and left out tool output, system notes and pasted content. That left 1,323 messages over 40 active days.

A regex pass sorted them into complaint categories. Then I read the 282 messages with a complaint signal by hand, because regex alone over-counts and misses sarcasm.

02 · Results

Complaints by type

  • Error I had to find myself3.7% → 2.4%38
  • "Show me it working"4.0% → 1.9%36
  • Unverified claim3.2% → 2.5%34
  • Too much text1.8% → 2.7%27
  • Generic "AI look" design2.3% → 2.1%26
  • Weak email draft1.4% → 2.7%26

Share of messages in July, then August 2026. Red means it got worse after the rule was written.

What I took from it

  • Unverified claims hurt most, even though they were not the most frequent. They are the only category that cost real money.
  • An automated check that passed was being used as proof of quality. It only proves what it measures. Now a screenshot has to be opened and described before anything is called done.
  • Two categories grew after I wrote the rule: long answers and weak emails. Writing a rule down is not the same as it being followed.

03 · Rules

Every failure became something that can be checked

  • Every number in an answer has its source next to it: a file path or the URL and what was seen in the live page.
  • Before showing a website: run the QA script, then open desktop and mobile screenshots and describe them.
  • A batch is done at N of N. If some items fail, say how many and why.
  • Answers have a hard length limit. Longer content goes into a file.