ENUA

Anastasiia Mudra · October 6, 2026

Trained to Please: Why LLMs Can't Score Things

Why LLMs are bad at scoring

The easiest way to get a score from a language model is to hand it some text and ask "how serious is this, from 1 to 10". The answer comes back fast, with an explanation, and it looks convincing. But a trivial issue and a real problem will get almost the same score, so that number cannot tell them apart.

  • LLMs score within a narrow range. Researchers find that LLM evaluators produce skewed score distributions (Stureborg et al., 2024) and compress the scale toward the middle: they inflate low scores and deflate high ones. In clinical scoring, even examples covering the whole scale in the request did not fix this (Zhang et al., 2026).
  • That is how they were trained. After pretraining, an LLM is fine-tuned on human ratings (Reinforcement Learning from Human Feedback): people compare answers, and the model learns to write the way people prefer. And people tend to approve of answers they agree with. This is how a model learns to adapt to whoever it is talking to and tell them what they want to hear (Sharma et al., 2023). Another study showed that this very tuning toward human expectations (alignment) makes a model pick the same scores again and again (Sato et al., 2026).

For a model trained to please, a "1" or a "10" is a harsh judgment and a potential conflict. So it picks the middle, or shifts everything into the "serious, but not a disaster" zone. As a result, both a trivial issue and something serious get roughly a 7, and the score no longer separates them.

What Jev is

In September 2026, TypeSafe AI released Jev. It is a model that does not generate text at all. It takes text or JSON plus predefined questions as input and returns typed answers:

  • Noul: the probability that the answer is "yes", from 0 to 1;
  • Choice: one option picked from a list, with a probability for each;
  • Score: a score on a scale you define: each level is described in words (for example, "1: cosmetic detail", "10: all data lost"), and the model picks the description that best fits what it was asked to score.

But the main difference between Jev and an LLM is how it was trained. During training, the model was rewarded not for answers people like, but for accuracy: if the model says it is 90% sure, it should be right about 9 times out of 10 (Reinforcement Learning for Calibrated Decisions). It does not explain or persuade: it returns a number that a program can act on directly.

That sounds like a fix for exactly the problem this article started with. It was worth checking on real tasks.

Two models, the same tasks

The comparison used a working product in which an LLM makes many small decisions every day: it answers "yes" or "no", picks an option from a list, and scores things from 1 to 10.

The LLMs used here were inexpensive but reasonably capable models, the kind usually picked for small decisions like these, rather than the most expensive ones on the market.

These tasks were given to both models at once: the LLM and Jev received the same data and answered independently. Every case where their answers differed was reviewed by hand to determine which one was right. Both models' scores were checked against the same scale, where every level has a clear description. The comparison deliberately used real production data rather than a specially prepared set: this shows how the models behave in everyday work, not on convenient examples.

Results

The first column of the table shows how often both models gave the same answer. The second shows which model turned out to be right after a manual check, in the cases where their answers differed.

Task typeHow often the answers matchedWho is right when the answers differ
Yes / no*96%Jev, in the cases that matter more
Choice from a list72%Jev in 20%, neither is better in the rest
Score on a 1–⁠10 scale**8%Jev in 73%, neither is better in the rest

* For a yes-or-no question, Jev returns not the answer itself but the probability that the answer is "yes", from 0 to 1. Here an answer counted as "yes" when that probability was 0.6 or higher.
** The low match rate here does not mean either model is bad: they simply gave different scores. The chart below shows how.

Yes / no answers: the LLM's mistakes cost more

In simple binary decisions there is almost no difference: 96% of answers match. However, the disagreements that do occur differ in nature:

  • Where the LLM was right, Jev mostly gave a probability close to one half, openly showing that it was unsure. Raising the threshold from the usual 0.5 to 0.6 removed some of these cases on its own.
  • Where Jev was right, the LLM's mistakes cost more: it let through what should have been rejected and rejected what mattered.

The LLM has no such probability: it answers "yes" or "no" with the same confidence, even when it is unsure.

Choice from a list: the LLM talks too much

Asked "which option", the LLM answers with a sentence, and the answer still has to be extracted from it. Returning exactly one option in the requested format is hard for it: part of its effort goes into explanation and formatting, and even a correct answer can get lost in that text. Sometimes it names several options at once, and then the wrong one ends up selected by accident. Jev picks from the list, so this kind of mistake is impossible for it by design. In the other cases where the models answered differently, both answers were correct: the models simply picked different but equally reasonable options.

Score from 1 to 10: half the scale does not exist for the LLM

Score from 1 to 10
The LLM pushes every score into the upper half of the scale, while Jev spreads them more evenly
LLM Jev
0%25%50%12345678910Share of scorestrivialcriticalScore from 1 to 10
The same data, scored by both models.

The data here is real, so the scores were never expected to cover the whole scale evenly: there were simply no catastrophic cases in it, which is why neither model gave extreme scores. What matters is not the distribution itself but the difference between the two models on the same data.

The LLM did not give a single score below 5: everything it saw seemed "fairly important" to it. Its average score was 6.8, and almost a third of cases got an 8. Jev uses the lower part of the scale and tells cases apart: a genuinely trivial issue gets a 2, exactly as the scale describes it.

A telling example: something that did not deserve attention at all was rated by the LLM as nearly critical, 8–⁠9 out of 10. Jev gave it a 4–⁠5.

The practical consequence: a threshold of "reject everything scored below 4" works with Jev and filters out trivial issues. With the LLM it would not have triggered even once.

While the LLM explains, Jev has already answered

Same decisions, same inputs
How much faster and cheaper Jev is than the LLM
LLM Jev
Response time
×20 faster
LLM5.20 s
Jev0.26 s
Cost per 1,000 decisions
×9 cheaper
LLM$0.490
Jev$0.054
Median time per answer and cost at provider prices. Shorter bar is better.
JevLLM
Median response time~0.26 s~5 s
Slowest 5% of responsesover 0.4 sover 29 s
Longest response3.7 s74 s
Cost per 1,000 decisions$0.054$0.49

The biggest gap is in 1-to-10 scores: here Jev is 35 times cheaper. The LLM attaches a lengthy explanation to every score, and every word of it has to be paid for. Jev returns just a number.

Jev has drawbacks too: it is only available as an external service, and sometimes it simply did not answer in time. In those cases the LLM made the decision, so a fallback is a must.

Conclusion

  • An LLM is a poor fit for scoring on a scale. It is trained to please, so its scores are compressed and tell nothing apart. A model trained for accuracy does noticeably better.
  • For yes/no decisions the quality is the same, but Jev gives a probability, and with it a threshold that can be tuned. It is also many times cheaper and faster.
  • When the answer is a choice from a list, text is unnecessary. The less text there is to parse, the fewer ways there are to go wrong.
  • The LLM stays where text is needed: finding a problem, explaining it, suggesting a fix. Jev does not generate text and does not try to.

So all of these decisions are now made by Jev, and it answers "yes" only when it is at least 60% sure.

Related reading