What we are building

Full benchmark suites such as lm-evaluation-harness take hours and real money to run. Most of the time the question is simpler: is this endpoint worth a closer look? This tool will point at any OpenAI-compatible endpoint, run a small curated set of prompts, and return a scorecard with quality, latency, and cost side by side.

Status

Coming soon. Not yet available.