Model bake-off bench for choosing open source AI models

By General Input

Stop picking models off leaderboards. Keep your own test set, shortlist candidates from Hugging Face, and run them side by side to see which one actually wins.

Integrations

  • Hugging Face
  • Google Sheets

Type

App

Categories

  • Engineering
  • Product

Build me a model bake-off bench I open every time we consider a new open model, so we stop picking models off leaderboards and start testing them on our own data. It has three views, Test Sets, Candidates and Results, plus a run history. The whole team works out of the same app: test sets are shared across the org, while each person keeps their own shortlist of candidate models.

The Test Sets view is where I keep my own evaluation material. A test set has a name, a description, and a list of test cases; each test case has a prompt, optional input text to go with it, and notes on what a good answer looks like. The test set also carries a rubric I write in my own words, for example correct SQL with no invented column names and under 200 words, and that rubric is what the grader uses later. Let me create, edit, duplicate and delete test sets and test cases inline, and store them in the app so the whole team sees the same ones.

The Candidates view is where I find models to try. Give me a search box backed by Hugging Face Quick Search for fast autocomplete as I type, and a fuller browse backed by Hugging Face List Models so I can filter by task, library and tags and sort by downloads, likes or most recently updated. Show each result as a card, and for every model I open, call Hugging Face Get Model to fill the card in with downloads, likes, license and last updated date. Put Add to shortlist and Remove from shortlist on each card. The shortlist is per person, so two of us can evaluate different sets of models against the same shared test set.

The Results view is a matrix: every test case in the selected test set is a row, and every shortlisted model is a column, so I can read the outputs side by side. Each cell shows that model's output, its score from the grader, and the latency and token counts for that call, and clicking a cell opens the full output plus the grader's written reasoning. Above the matrix, show a per model summary row with average score, average latency, total tokens, and the recommended pick called out.

The big action is a Run bake-off button that kicks off a background agent. The agent runs every test case in the selected test set against every model on my shortlist using Hugging Face Chat Completion, capturing the output, the latency, and the prompt and completion token counts for each call. It then grades each output against the rubric I wrote for that test set, producing a score and a short written reason for that score. Finally it writes the per model scores, the reasoning behind each score, and a recommended pick with a sentence explaining why, back into the app, so the Results view fills in as the run progresses. Show progress while it runs, and keep partial results if one model fails rather than losing the whole run.

When calling Chat Completion, the model id can be suffixed to control routing, for example :fastest, :cheapest, :preferred, or an explicit provider name. Let me choose the routing policy for a bake-off so speed and cost sit next to quality, and record which provider actually served each call alongside the latency and token counts.

Keep every past bake-off in a run history, stamped with the date, the test set used, the models compared, the routing policy and the winner. When a new model drops I want to open a past run, add the new model as a candidate, re-run the same test set, and see how it did against the model we picked last time, so give me a compare view that puts a new run next to an earlier one and highlights where scores moved.

Add an Export scorecard button that appends the finished comparison to a Google Sheet using Google Sheets Append Values, one row per model per run carrying the run date, test set name, model id, average score, average latency, total tokens, license and whether it was the recommended pick, so I can share it with people who will not open the app. Let me pick the destination spreadsheet and tab.

Store test sets, shortlists, runs and their results in the app so nothing is lost between sessions. Test sets and run history are shared across the team, and shortlists are per person.

Related prompts

Explore more prompts
Call overdue Xero customers with an AI collections agentLocal listing health board for every location you manageLet support send one-off Loops emails without an engineerStop cold emails to anyone with a live deal in PipedriveiMessage campaign console with pre-flight checks and delivery boardLinkedIn Ads budget pacing dashboard for every client accountFront desk appointment confirmation board for the next 3 daysGive your team Looker numbers without buying more seatsBuild audience segments from product usage and push to LoopsTurn the people who engage with your posts into Pipedrive leads