Everything else in this cluster is about renting intelligence. This page is about the other option: models you download and run yourself, on hardware you control. Five years ago that was a research hobby; today the open-weight world is a parallel ecosystem serious enough that for some workloads it is the rational default, and for a subset of regulated work it is the only comfortable answer.

The families that matter

The open-weight field consolidates around a handful of well-backed families, each with a model range from small to large:

  • Llama (Meta): the most widely deployed family and the ecosystem anchor, with the broadest tooling support. Custom licence; permissive for normal businesses, with conditions worth one careful read. Covered in more depth in the Meta guide.
  • Mistral (Mistral AI, France): the European champion, popular where EU headquarters and data posture matter, with genuinely permissive licensing on many models and strong small-model quality.
  • DeepSeek (China): the family that shocked the market in early 2025 by matching far more expensive models at a fraction of the training cost, published under permissive terms. Capability per pound is the draw; for some firms, provenance and their own risk committee’s view of it will be part of the assessment.
  • Qwen (Alibaba): a broad, fast-improving family with strong multilingual performance, widely used and largely permissively licensed.
  • Gemma (Google): Google’s open family, smaller than its Gemini flagships, well-documented and convenient for teams already in Google’s ecosystem.

The shortlist at a glance:

FamilyBackerLicence postureKnown for
LlamaMetaCustom community licenceWidest deployment, deepest tooling
MistralMistral AI (France)Largely Apache 2.0Small-model quality, EU comfort
DeepSeekDeepSeek (China)MIT-styleCapability per pound
QwenAlibabaLargely Apache 2.0Multilingual breadth
GemmaGoogleOpen weights, Google termsDocumentation, ecosystem fit

Names and version numbers churn constantly; the families and their characters persist. Anything specific this page could claim about “the current best” would be stale within a quarter, which is itself a lesson for how firms should decide.

Three reasons to run your own

Data control. Prompts, documents and outputs that never leave infrastructure you own. For workloads where sending client data to any third party is unacceptable or just expensively awkward to paper, this is decisive, and it is the reason the option exists on this site’s own Software Factory in the first place.

Cost shape at volume. APIs price per use; your own hardware prices per month regardless of use. High, steady volume eventually crosses the line where owning is cheaper. The calculation is unglamorous: expected tokens, hardware rental, an allowance for the person who keeps it running. Do it before believing either answer.

Independence. No provider can reprice, rate-limit, retire or change the model underneath you. For systems meant to run for years, that stability has real value, and it compounds with the switchability argument made across this whole cluster.

Against all three sits the cost nobody advertises: you become the host. Updates, security, monitoring, capacity, backups. It is the same discipline any production system needs, which is fine if you have that discipline and a liability if you do not.

The middle path: hosted open models

Between “pay OpenAI per token” and “rack your own servers” sits the option that suits most firms flirting with open weights: hosted open models. Major clouds and specialist providers serve Llama, Mistral, DeepSeek and Qwen per use, sometimes at striking prices. You get the model diversity and some of the independence without becoming a hosting operation. It is also the cheap way to evaluate whether an open model handles your workload before any hardware conversation happens.

The checklist

  • Licence read and cleared for your intended use
  • Workload benchmarked on your own tasks against a hosted version first
  • Volume economics calculated, not assumed, with the crossover point written down
  • Hosting owner named, with monitoring, updates and backups in the plan, not the appendix
  • Data protection assessed as for any processor: self-hosting moves the obligations, it does not remove them (governance guides, DPIA starter)

Where I fit in

This is home ground: I run models and AI systems on my own infrastructure every day, and the Software Factory explicitly offers both homes, my managed infrastructure or a cloud server in your own account, hardened before any business data enters and watched for as long as it runs. If you suspect a workload of yours belongs on open weights, bring the volume numbers to a call; the economics usually make the decision for us within the hour.