What You're Actually Buying When You Pick an LLM Vendor
I started this analysis with a practical question: once the leading models score within a few benchm 2026-7-23 07:6:18 Author: hackernoon.com(查看原文) 阅读量:2 收藏

I started this analysis with a practical question: once the leading models score within a few benchmark points of each other, what should companies actually compare when choosing an LLM provider? I compared 33 models from 15 providers to find out, and the number everyone leads with turned out to be far less useful than the differences behind it.

Vendor comparison pages still lead with one figure: a reasoning benchmark score. But among the leading models, that number is no longer enough to make a vendor decision. The top ten now sit within a six-point spread on GPQA Diamond, the PhD-level science benchmark most vendors quote. Claude Mythos 5 leads at 94.4%, while the tenth-ranked model is at 89%. A gap that small doesn't tell you which vendor to pick.

Among leading models, intelligence alone is no longer the main differentiator. Vendors now differ in three important ways: how often the model refuses to answer, how much control you get for the price, and, for some providers, which national content rules shaped the model. None of these show up on the comparison page. All three show up in how the model behaves in production. Choosing the wrong model can mean higher costs, unexpected refusals in production, vendor-driven behavior changes, or additional work to correct jurisdiction-specific bias.

Reasoning has converged

Until recently, a five- to ten-point gap in reasoning scores was a meaningful signal. Today, Gemini 3.1 Pro, Fable 5, Opus 4.8, Qwen 3.7 Max, and GPT-5.4 are within a few points of each other on GPQA.

Model

Provider

GPQA Diamond

Access

Claude Mythos 5

Anthropic

94.4%

Closed

Gemini 3.1 Pro

Google

94.3%

Closed

Claude Fable 5

Anthropic

94.0%

Closed

Claude Opus 4.8

Anthropic

93.6%

Closed

Qwen 3.7 Max

Alibaba

92.4%

Closed

GPT-5.4

OpenAI

92.0%

Closed

GLM-5.2

Zhipu

91.2%

Open

DeepSeek V4 Pro

DeepSeek

90.1%

Open

Harder tests like Humanity's Last Exam still show real separation at the top, and this is where the frontier differentiates once GPQA stops clearly separating the leaders.

Claude Mythos 5 leads at 64.5%, Claude Fable 5 follows at 53.3%, Opus 4.8 sits at 45.7%, and Gemini 3.1 Pro is close behind at 44.4%. Grok 4.20 posts 50.7%, notably ahead of both Anthropic's Opus 4.8 and Google's flagship on this specific test, despite trailing on GPQA.

Most procurement conversations don't get this far. They stop at the headline number, which by itself no longer distinguishes a serious vendor from another serious vendor.

Refusal rates vary by provider, not just by model

The labs with the strongest reasoning scores also sit highest on guardrails. That's a design choice, not a byproduct of capability. RefusalBench, a May 2026 study using biology-adjacent research prompts, found that Anthropic's models refuse roughly 21 times more often than the market median, with refusal rates across all tested models ranging from 0.1% to 94.6%. Provider identity was a strong predictor of where a model landed on that range.

For companies, this means a closed frontier model comes with a fixed refusal policy you don't control and can't adjust through configuration. For a regulated product or anything customer-facing, that predictability is often worth paying for. For an internal research pipeline or an agent that needs to work through edge cases, the same policy becomes friction with no override.

The strongest alternatives offer more control at a lower cost

A separate cluster of models delivers near-frontier reasoning with far fewer refusals: Grok 4.20 and 4.3, GLM-5.2, DeepSeek V4 Pro, Qwen 3.7 Max, and Kimi K2.6.

That group mixes closed models with a lighter refusal policy (Grok, Qwen 3.7 Max) and genuinely open-weight models (GLM-5.2, DeepSeek V4 Pro, Kimi K2.6) - worth keeping straight, since only the open-weight ones can actually be fine-tuned or self-hosted.

The strongest models in this group trail the top closed frontier by only two to four GPQA points, while others trade more reasoning performance for lower cost and greater control.

Three models in the broader open-weight cluster: GLM-5.2, DeepSeek V4 Pro, and Qwen3.5 397B - combine strong reasoning, open weights, and low guardrails at the same time.

Here's what that gap looks like model by model, measured directly against Fable 5, the current top performer:

Model

Access

GPQA

Guardrail

GPQA gap vs. Fable 5

Claude Fable 5

Closed

94.0%

9.0

Qwen 3.7 Max

Closed

92.4%

4.5

-1.6

GLM-5.2

Open

91.2%

3.5

-2.8

DeepSeek V4 Pro

Open

90.1%

3.5

-3.9

Grok 4.20

Closed

88.5%

3.0

-5.5

Qwen3.6 35B

Open

86.0%

3.5

-8.0

Mistral Large 3

Open

72.0%

3.0

-22.0

Hermes 4 405B

Open

70.5%

1.5

-23.5

The takeaway from that table: several models stay within four GPQA points of Fable 5 while offering lower guardrails.

The open-weight models in this group: GLM-5.2, DeepSeek V4 Pro, also let teams modify refusal behavior and control deployment directly, since fine-tuning and self-hosting are only available where the weights are open.

Below Qwen3.6 35B, the gap widens fast - those models are trading real reasoning capability for a clean, low-guardrail base rather than staying competitive on raw intelligence.

The trade-off is straightforward. A closed model's refusal policy is fixed for as long as you use the API, and it can change without notice when the provider updates it. An open model's refusal behavior is a starting point that can be fine-tuned or removed.

The gap between closed and open frontier performance has narrowed to roughly six points this year. In practice, open weights are no longer the fallback option for teams that can't afford closed models, they're a deliberate choice for teams that want to own the behavior of the system they're running.

Matvii Diadkov's image-90f0d8

The scatter plot maps the models into three broad clusters: closed frontier models at 93–94% GPQA with guardrail scores of 6.5–9.0, an unlocked-and-smart cluster at 88–92% GPQA with guardrail scores of 3.0–4.5, and a maximally-unlocked cluster (Hermes 4, Mistral Large 3) at 70–72% GPQA with guardrail scores below 3.0.

The most commercially relevant area is the lower-right of that map: models that stay close to the reasoning frontier while giving teams more control over cost and behavior.

The other areas of the map cover different needs. Gemma, Cohere's Command A, Phi-4, and Claude Haiku 4.5 are weaker on reasoning and still heavily guardrailed - these are built for edge and on-device deployments, enterprise RAG with grounded citations, and cheap high-volume inference, not general-purpose reasoning.

Hermes 4 and Mistral Large 3, in the maximally-unlocked cluster above, are weaker on reasoning but almost entirely unlocked - their value isn't raw capability, it's serving as a clean base for a team that wants to fine-tune its own aligned model rather than inherit anyone else's policy.

A regional filter applies to Chinese open models

The major Chinese open-weight models: DeepSeek, Qwen, GLM, Kimi, MiniMax, include a political filter that Western models don't have. This follows directly from China's Interim Measures for the Management of Generative AI Services, which require public-facing models to avoid content that undermines social stability and require domestic labs to pass a government security review before release. The behavior is a licensing condition, not a company preference.

Every model refuses something. Ask a Western open model like Llama for its most-refused topics and you get universal categories: child exploitation, hate speech, violence. Chinese models refuse those too. What they add is a second, regional layer specific to China: Tiananmen 1989, Taiwan's status, Xinjiang, criticism of Xi Jinping, the Hong Kong protests, and even politically sensitive references such as Winnie the Pooh. Northeastern University's "R1dacted" study formalized exactly this split between global and local refusal categories.

The question isn't whether a model censors anything - every model does - it's which additional topics get filtered, and whether those topics intersect with your product.

It operates in two layers:

  • The first is a live filter on the hosted API that can cut off an answer mid-sentence on topics like Tiananmen Square; this layer disappears entirely once you self-host the open weights.
  • The second is trained into the model during fine-tuning and survives self-hosting. Ask the same model to describe conditions in Xinjiang, for instance, and a Chinese model will typically default to language about social stability and ethnic harmony, where a Western model will reference documented allegations directly, that's the weight-level layer, not the API filter, and self-hosting doesn't touch it.

Enkrypt AI's testing found that even after standard jailbreak attempts, roughly 91% of DeepSeek R1's answers on China-related political topics still leaned pro-Beijing. Self-hosting gives you data control. It does not remove the underlying bias, which requires a separate de-censoring fine-tune.

In practice, this filtering is unlikely to affect most coding, data analysis, agent, and standard business workloads. It matters specifically for products touching geopolitics, Taiwan, Xinjiang, or Hong Kong. For those products, the filter is a cost to plan for, either through a fine-tune or by routing that traffic to a model without the same China-specific alignment. It is not a reason to avoid otherwise strong, low-cost models.

What this means for vendor selection

Once benchmark scores are treated as roughly equivalent at the top, the decision comes down to three variables that trade off differently depending on the workload: how much refusal policy you're willing to accept as fixed, how much you're paying for capability versus control, and whether jurisdiction-specific behavior applies to your use case.

A regulated or customer-facing product typically favors locking down the first variable and accepting the cost of a closed frontier model. An internal tool or agent framework typically favors the second, where cost differences of 5x to 100x and full policy control matter more than a few points of benchmark score. A product with no geopolitical exposure can set the third variable aside; one that has such exposure needs to budget for it directly.

Routing by task, not by vendor

Standardizing on a single vendor across every use case ignores that these three variables don't move together. A more workable approach routes by task: closed frontier models for the highest-stakes reasoning and customer-facing work where predictable behavior is part of the value; the open, unlocked cluster for internal infrastructure, agents, and cost-sensitive volume work; and neutral open bases like Hermes 4 or Mistral Large 3 as a starting point for teams building their own fine-tune.

The rankings will change quickly as new models arrive. The decision framework is less likely to change: capability, refusal policy, control, cost, and jurisdiction still need to be evaluated together.

My conclusion is simple: companies should stop treating LLM selection as a one-vendor decision. Benchmark scores will keep converging, while refusal policies, costs, control, and jurisdiction will not. The practical answer is to route each workload to the model that fits it best.


文章来源: https://hackernoon.com/what-youre-actually-buying-when-you-pick-an-llm-vendor?source=rss
如有侵权请联系:admin#unsafe.sh