Skip to main content
AI Models

Qwen, GLM, Kimi, MiniMax: How Open-Weight Models Changed Chatbot Pricing

Cortexvia Research25 Aug 20264 min read

Two years ago, running a capable language model meant paying one of three labs per token. In 2026, several labs publish open-weight models, meaning the model files themselves are released and anyone can host them: Alibaba's Qwen family, Zhipu AI's GLM, Moonshot AI's Kimi, MiniMax, alongside releases from Meta, Mistral, and others. Their quality has climbed to the point where they handle most website chatbot questions well, and hosting competition has driven the price of running them down sharply.

That shift is the reason chatbot pricing looks the way it does now. Here is what changed, and what it means when you are choosing a plan.

What open weights actually change

Anyone can host them. Dozens of inference providers run the same open models and compete on price and speed. Closed models are only available from their lab, at that lab's price.

Prices fell. With many providers running equivalent models, the cost per million tokens for a capable model fell to a small fraction of what closed frontier models charge. For a chatbot answering short questions from retrieved passages, that cost is now a rounding error per conversation.

The quality floor rose. The gap between a good open-weight model and a closed frontier model, on the grounded single-fact questions a support widget sees, is small. The gap is larger on hard reasoning, which is why frontier models still have a place on deep tiers.

What that did to chatbot plans

Flat plans with generous message allowances became normal. When each answer cost real money, vendors metered conversations tightly and charged overage. When answers cost fractions of a cent, a daily allowance on a flat monthly plan is easy to offer and easy to budget. That is the model Cortexvia uses; see the pricing page and how much does an AI chatbot cost.

Tiers replaced vendor names. Platforms package models into tiers by speed and depth rather than making buyers pick a lab, because the lab matters less than it did and changes often. Cortexvia's tiers are described by purpose on the models page.

Free plans got real. A free tier that lets you run a live chatbot on a live site, rather than a fourteen-day trial, is affordable for vendors because of these costs. The free plan is an example.

The vendor lock-in question

Buyers used to ask "what if the model vendor raises prices or shuts the model down." With open-weight models in the mix, a platform can move between hosts or models without the customer noticing. That is only true if your content is portable, which it is when the platform works from your own documents and pages rather than from custom training. Ask any vendor whether you can export your sources and conversations; the honest ones say yes without hesitation.

Where closed frontier models still win

  • Hard, multi-step reasoning across long documents
  • Following complex instructions with many exceptions
  • Less common languages, where the largest models are still measurably better

For a website chatbot, those show up on the deep tier and on internal documentation bots, not on the widget answering "what are your hours". The reasoning models guide covers when to pay for depth.

What to check as a buyer

  1. Is the price flat, and what is the message allowance? Overage terms tell you whether the vendor is still paying frontier prices per token.
  2. Are tiers described by what they do, and can you switch between them?
  3. Can you export your sources and conversations?
  4. Is the price shown the price charged in your currency?

Those four questions tell you more about a vendor's cost structure than any model name on the pricing page.

Frequently asked questions

Are open-weight models safe to use for customer-facing bots? Yes, when the platform grounds them in your content and adds a handoff. Safety on a support widget comes from grounding and instructions, not from which lab trained the weights.

Is "open-weight" the same as "open source"? Not quite. Open-weight means the model files are released under a licence that permits hosting and often commercial use; the training data and full recipe usually are not. For a buyer the practical effect is the same: many providers, competitive prices.

Does Cortexvia run open-weight models? Cortexvia exposes tiers rather than model vendors, and the tiers are chosen for speed, accuracy, and cost on grounded questions. The tier test in the model choice guide is the way to pick one.

Ready to Get Started?

Build your chatbot in minutes with Cortexvia. Free tier available.