Luke Angel
← back to the journal
Tag

#testing

9 entries
method
Twenty-Two Behavioural Criteria: What Replaced the Structural Grader My AI Agents Kept Fooling
“Scoring went from 28 seconds to 2, and got strictly harder to fool.”
Aug 06
method
The AI Grader That Passed the Same Bug Three Times: 229 Checks, Three Agent Workspaces, One Blind Spot
“A perfect score, three times, over the same bug. The checks weren't wrong. They were pointed at the wrong thing.”
Aug 04
method
A Scorecard, Not a Vibe: What I'd Need Before an AI Coding Agent Touches My Codebase
“An agent asked to judge its own completion will do what anything does in that position.”
Jul 14
projects
The Battery Test: Runtime Is a Dial, Not a Number
“Runtime isn't a spec you read off the battery — it's a dial you set in firmware.”
Jun 21
projects
The City Walk: Path Loss ~4 in Dense Urban — and the Cache That Wouldn't Die
“The chip was telling the truth the whole time. The firmware just wouldn't repeat it.”
Jun 04
projects
The Range Test: 1,250 Feet Direct — and Miles Through the Mesh
“Zero-hop is the floor, not the ceiling. The direct radio cleared the bar on its own — and the mesh I didn't build was already turning feet into miles.”
May 28
projects
A Scorecard, Not a Vibe: How I'll Decide Whether to Build the Collar
“A dot moving on a map isn't proof. A dot that moves reliably at 300 meters, for a week on a charge, is.”
May 26
craft
Chaos-pass replaces tests-pass
“Steady-state-only passing is insufficient. The cluster has to survive being broken.”
May 15
tools
Prompt regression — the rough first version of our test set
“We're testing prompts the same way we test code. We just haven't earned the vocabulary yet.”
Nov 15