local-first-model-router

Experimental: local-first inference governed by explicit policy, measured capability, and strict tool authority.

incident-response-agent

Proof of concept: human-approved incident remediation with bounded authority and an observable execution trail.

agent-cockpit

Prototype: keeps coding-agent dispatch decisions in versioned repository state instead of chat history.

llm-dyno

Pre-release: measures deployed model behavior with deterministic grading instead of treating a model name as a benchmark.

structured-output-agent

Experimental: treats structured output and tool use as validation boundaries, not prompting conventions.

react-planning-agent

Experimental: makes the ReAct loop typed, bounded, and observable, including how it stops.

toolcall-repair-bench

Concluded experiment: follow-up tests showed the apparent tool-call problem was mostly a serving-stack artifact.

local-tool-proxy

Concluded prototype: worked on a narrow raw-output path; later evidence showed modern parsed endpoints usually do not need it.