Project Active

Local AI Harness

A controlled environment for testing whether local AI models can make real code changes, with the harness, not the model, owning verification.

Type
Project
Status
Active
Started
2026-08-15
Themes
local-ai, evaluation, engineering-workflows

Why it exists

Can an open model running on a personal computer make useful code changes reliably enough to become part of a real workflow? Model quality, inference cost, hardware limits and workflow reliability all decide that, and I wanted to answer it through real work rather than leaderboards. The premise: you can't trust a model to enforce its own rules, so invert control. The model proposes. The harness validates, executes, reviews and records what actually happened.

Approach

The harness. It gives a model a bounded task in an isolated workspace, controls the available tools, records actions, runs checks and produces evidence for review. Implementation, verification and human acceptance are separate steps. It has been exercised against a throwaway project, a real defect in my Recall app, and the harness itself.

A plan that survived an audit. I audited my first build plan, found eleven defects, and wrote a phased final version with strict phase boundaries, tiered success criteria and an explicit cut line. It reuses existing guardrail libraries instead of rebuilding them, guards against the runtime silently truncating context below the configured window, uses a separate smaller model as reviewer, and includes a Windows path-escape test suite (UNC paths, device namespaces, alternate data streams, short names, junctions). A guard only counts once I have seen it fire.

Hardware qualification. The workstation is an RTX 3080 with 10GB of VRAM and 128GB of system memory. An audit found all four RAM sticks running at DDR4-2133 instead of their rated 3200, roughly a 50% memory-bandwidth deficit, because XMP had never been enabled. With 10GB of VRAM, dense 70B models aren't practical, so the plan centers on mixture-of-experts models. I measure context size, memory use, GPU residency, prompt processing, generation speed and errors.

A useful failure. In one task on Recall, a local 30B worker made a single edit and reran the suite to 83 passing tests, staying within its allowed boundaries. An engineering review still found an atomicity and race concern in the fix. Other trials stopped before making the change, printed text that looked like a tool call without executing it, or claimed work that never happened.

What I learned

Passing tests is evidence, not a verdict. Reliability comes from the whole system around a model: tools, instructions, context, recovery, verification and human judgment. Those failures defined what the harness has to verify independently.