Status: Pre-release
Code: github.com/cameronqj/llm-dyno
A model name is not a performance result. Runtime, quantization, hardware, concurrency, context length, and prompting scaffold all change what a deployment can actually do.
llm-dyno measures the endpoint in front of you. It uses one OpenAI-compatible interface, deterministic code-based graders, context and concurrency tests, and math/options suites designed to separate capable deployments instead of producing easy perfect scores. Prompt scaffolding is recorded as an experimental variable rather than disappearing into the result.
Three private targets have completed the full artifact pipeline and produced validated profiles and self-contained reports. Their identities and measurements stay local by design, so the public repository does not contain a model leaderboard or a comparative performance claim. The harness is still pre-release; the evidence currently supports the measurement pipeline, not a universal ranking.