Trained to Please: Why LLMs Can't Score Things
Language models avoid harsh scores because they are trained to please. A comparison of an LLM and Jev, a model that returns decisions instead of text.
· 9 min read · llm, ai-agents, engineering, benchmarking