Odd tests.
Useful signals.
The open bazaar for community-made LLM benchmarks. Publish the method, keep the official test set sealed, and bring receipts.
Open the mystery cratePick an aisle
What deserves a reality check?
Editorial categories, with plain-language labels for every market sign.
Fresh stock
New on the shelves
Recently published benchmarks, ordered by publish date—not pageviews.
Curator’s cart
Worth a closer look
A small, hand-picked mix from across the bazaar.
Receipts just in
Claims you can inspect
Fresh model results with exact versions and provenance attached.
Best sellers
Busy checkout counters
The most distinct valid model runs in this preview—not the most pageviews.
Open method, sealed official set
Show your work without giving away the test.
Benchmark pages publish the purpose, public examples, scoring recipe, limitations, and result receipts. Official scored questions stay with a controlled runner instead of becoming an easy training-data download.
A model service may still observe evaluation prompts sent to it. Sealed reduces casual exposure; it does not mean impossible to leak.
Read the trust modelPurpose, samples, scorer, limitations.
Author-controlled evaluation, away from the browser.
Exact version, model, track, metrics, and evidence.
Got a strangely useful test?
Set up your stall.
A benchmark idea does not need a paper, package, or giant test suite to be useful. Start with the question it answers and the limits it has.
Publish a benchmark