LAB NOTEBOOK · PLANNED
These are experiment directions — not released products and not verified production capabilities. They are published because the lab’s process is public: you should be able to see what we intend to test, before we test it.
Bounded MCP service
Prototype one small, auditable MR BIG capability with FastMCP and measure integration effort, permissions, and client compatibility. The question: how much discipline does it take to expose exactly one bounded capability — and nothing else — through MCP?
Repository context quality
Compare graph-backed code context with an existing workflow, tracking retrieval recall, stale-index behavior, and review usefulness. The question: does persistent codebase context actually improve review, or just feel impressive?
Agent-assisted security review
Evaluate Deepsec in a sandbox alongside static analysis and human review, measuring validated findings and false positives. The question: can a coding-agent approach to vulnerability discovery earn a place in a serious review pipeline — sandboxed, validated, and honest about false positives?
Terminal coding agent
Evaluate Qwen Code on a non-sensitive test repository using a governed model configuration and explicit tool-safety boundaries. The question: what does a governed, inspectable terminal agent look like when the safety rails are configured by someone who means it?
EXPERIMENT STANDARD
Define the question, isolate the system, record the inputs, preserve the evidence, and publish limitations alongside results. Experiments are clearly labeled until results are verified.
Each direction ties back to the public projects: the bounded MCP service to MR BIG, the context and security evaluations to the Buzz Desk’s reporting, and the terminal-agent work to the reliability program. Results will be published as Lab Reports when verified.

