Category
Method.
How I work — prompts, evals, the four columns.
9 entries
method
Forty-Four Harness Bugs, Zero Local LLM Limitations: an Accounting
“A model that looks broken is a harness bug until proven otherwise.”
Aug 19
method
What a Bug Fix Costs on Two DGX Sparks: 16 Concurrent AI Agents, 4.45× Cheaper Than Serial
“The ceiling I'd been running against was a benchmark's sample size, not a limit of the machine.”
Aug 14
method
Nine Lines of Verification That Beat a Six-Agent AI Swarm
“It tried to stop 52 times. It was refused 51. That is the whole mechanism.”
Aug 12
method
1,200 Lines of Multi-Agent Orchestration, Beaten by One Local LLM Agent
“The orchestration didn't lose because multi-agent is a bad idea. It lost because it was insurance against a blindness that no longer existed.”
Aug 10
method
Twenty-Two Behavioural Criteria: What Replaced the Structural Grader My AI Agents Kept Fooling
“Scoring went from 28 seconds to 2, and got strictly harder to fool.”
Aug 06
method
The AI Grader That Passed the Same Bug Three Times: 229 Checks, Three Agent Workspaces, One Blind Spot
“A perfect score, three times, over the same bug. The checks weren't wrong. They were pointed at the wrong thing.”
Aug 04
method
Sixteen Requirements for an Agentic Coding Swarm, All of Them Scar Tissue
“Last-writer-wins is invisible to any score-based gate. The work vanishes and the number goes up.”
Jul 21
method
A Scorecard, Not a Vibe: What I'd Need Before an AI Coding Agent Touches My Codebase
“An agent asked to judge its own completion will do what anything does in that position.”
Jul 14
method
The Local LLM Bill: What a Bug Fix Has to Cost Before Two DGX Sparks Make Sense
“Per-token pricing prices the tokens you burn, and search burns tokens by design.”
Jul 07