The fkra journal
· 3 min· 93 reads

A better model will not save you from a bad rubric

Roboflow's write-up of GPT-5.6's vision performance is a good result and a reminder: "best" is only meaningful against a task you have defined.

Roboflow published a piece this week arguing that GPT-5.6 Sol is the strongest vision model OpenAI has shipped. Their evidence is specific, their test set is public, and that already puts the claim ahead of most of what gets posted about model quality.

It is a good result. It is also, read too quickly, a procurement decision heading for trouble.

Because the word doing all the work in that sentence is "strongest".

Vision is not one thing

Reading a dense table. Counting objects in a crowded frame. Spotting a small defect on a large surface. Describing a scene.

Four tasks, four different failure modes, one word covering all of them. A model can move decisively ahead on three and stay flat on the fourth — and every summary of that situation is a choice about which one the summariser happened to care about.

Usually that choice is invisible, including to the person making it.

This is not a new problem and it is not really about models. "Strong analyst" is not a measurement either. It is a compression of several judgements against an implicit standard, and the standard is where all the disagreement actually lives. We just do not notice, because with people we at least argue about the standard out loud.

Write the rubric first

The discipline that helps here is unglamorous and slightly annoying: define what a good answer looks like before you see any answers.

Rubrics written after the outputs arrive tend to describe the outputs. Not through dishonesty — through the ordinary human tendency to find the criteria that the impressive thing satisfies. You will do it, I will do it, and neither of us will notice.

A rubric written in advance also survives the model changing underneath you, which the leaderboard position will not.

Score against cost per sample
Illustrative, not measured. Note the log axis: cheapest to dearest spans three orders of magnitude. The scores do not.

The chart above is invented, but the shape is the argument: the spread on cost is enormous and the spread on score, for one particular task, often is not. Which means the interesting question is never "which is best" — it is "which is good enough here, and what does the gap cost me per thousand calls".

Where a public benchmark earns its keep

Aggregate benchmarks are excellent at exclusion. They tell you a model is not in contention, and in a field shipping weekly, cheap exclusion is worth real money.

They are poor at selection, because selection depends on a task distribution the benchmark has never seen. Roughly:

Question

Public benchmark

Your fifty examples

Is this model in the running?

Good

Overkill

Will it handle my documents?

No signal

The whole point

Will it still be right in six months?

Expires

Re-runnable

Can I defend the choice?

Appeal to authority

Evidence

Fifty examples

So use the leaderboard to build a shortlist, then decide with your own evaluation.

Concretely: assemble fifty examples from your real workload — including, and this is the part people skip, the ones your current process gets wrong. Score them against the definition you wrote first.

Fifty is small enough to build in an afternoon and large enough to embarrass a confident assumption. That set will outlive several model generations, which is more than most benchmarks manage, and it turns "which model is best" into a question you can actually answer for yourself.

Roboflow's write-up is worth reading — and then worth ignoring in favour of your own fifty.

0
93 views
evaluationrubricsbenchmarks
MA
mosab alrasheed
Get the next entry

One email when we publish. Research, product decisions, and what teams report back.