Design Status: Released v0.1.0

Model Selector

Match a model to a workload using stated constraints instead of vibes or leaderboard rank.

Runs in your browser. Nothing you enter is sent anywhere.

Note

How this catalog dates its data, and why a stale entry is not just an old price

Every candidate below carries its own price effective date and source. On page load, this tool compares each date against today in your own browser and flags anything older than 90 days, about one quarter. That threshold is a judgment call, stated plainly rather than hidden. It is roughly how often a major vendor changes list pricing or ships a new flagship model, so past it a comparison risks missing a change that would have altered the outcome.

This tool ships no capability, latency, or throughput rating of its own for any candidate. Those three axes are modeled only once you rate a candidate on them, from your own evals or experience, using the controls on each ranked row. An axis with no rating is named plainly rather than silently dropped, and a rated axis is marked rated by you rather than shown as though it were catalog fact, in both the ranked list and the export.

Workload requirements

Compared against whatever capability rating you supply per candidate below. This tool ships no capability rating of its own.

Tokens the longest single request needs, prompt plus expected output.

Milliseconds tolerable before a response feels broken. Leave blank for no ceiling, for example batch or async work.

Maximum dollars per million tokens accepted, blended across input and output. Leave blank for no ceiling.

Ranking weights

These weights only affect ranking among candidates that already pass every hard constraint above. They never change which candidates are eliminated. Raise a weight to care more about that axis, relative to the others.

Cost assumption

Every value below is stated, not hidden.

Input token share
75% from default assumption
Output token share
25% from default assumption

Ranked candidates

Candidates that pass every hard constraint, ordered by the weighted score above.

Ready

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Disqualified candidates

Eliminated by an objective hard constraint, never by an editorial rating.

Ready

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Tradeoff, the runner up

What follows from the stated weights, and what would have to change for the runner up to win instead.

Ready

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Unanswered questions

Gaps this tool noticed in the stated requirements.

Ready

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

    Recommended evaluation plan

    What to do before trusting this shortlist with real traffic.

    Ready

    Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.
      Export
      Note Why does this not just tell me the best model

      There is no universal best model. There is only a model that fits the constraints stated above, weighted the way they are weighted above. Change the accuracy bar, the cost ceiling, or any weight, and a different candidate can win, honestly, because the inputs changed.

      The tradeoff panel names the runner up and states exactly what weight would have to move, and by how much, for it to outrank the top candidate. That is the whole mechanism. Nothing here consults a live benchmark or a model provider.