LAB REPORTS
Results, released.
Reports, experiment records, and technical writing from the lab — evidence first, claims labeled, limitations published. Every release carries primary links, confidence levels, and practical limits.
-
PROJECT: BIGSTEVE — One Spec, Many Builders
The new BigSteveLabs.com, rebuilt as an experiment: multiple independent AI builders, one frozen specification, one scorecard. This site is one of the builds.
→
-
Lab Notebook: Current Experiment Directions
Four planned experiments — a bounded MCP service, repository context quality, agent-assisted security review, and a governed terminal coding agent. Directions, not claims.
→
-
Why We Break Things On Purpose: Deterministic Failure Testing
The case for deterministic failure-and-recovery testing in AI agent systems — and why reproducible failure beats ‘we tested it.’
→
-
Agent Rescue: Anatomy of a Failure at 73%
An AI job failed at 73% on an HTTP 429. The lab demonstrated the full recovery: identification, human-approved provider switch, preserved progress, resumption to completion.
→
-
Open Source & Developer Report #001 — RC1
FastMCP, Qwen Code, code-review-graph, pi, and Deepsec — checked for recent activity, licensing, adoption signals, concerns, and practical suitability. Corrected current edition; window: July 21-28, 2026.
→
