2. Key Players and Models

Introduction

Now that we’ve explored the foundational architecture of large language models (LLMs) and today’s AI landscape, let’s map the ecosystem of key players and models shaping this transformative technology. The field has evolved rapidly – new providers have emerged, reasoning has become an adjustable capability rather than a separate product line, and the open-weight ecosystem has repeatedly shifted the balance of power. Whether you’re considering commercial solutions, open-source options, or local deployments, understanding the ecosystem is essential for selecting the right tool for your needs.

What will I get out of this?

By the end of this section, you will be able to:

  1. Differentiate between major commercial and open-source LLM providers, including their key offerings and strengths.
  2. Compare foundation models, fine-tuned models, and specialized models, understanding their characteristics and use cases.
  3. Explain the lifecycle of an LLM, including the training process, fine-tuning methodology, and inference.
  4. Explain extended reasoning – what a model does differently when it deliberates, and when the extra latency and cost are worth paying for.
  5. Place a model on the size spectrum, from on-device small language models through to frontier systems, and recognize which tier a workload calls for.
  6. Analyze the trade-offs between model types and licensing postures, including what building on a given vendor’s open weights commits you to. (Turning these trade-offs into a deployment decision is the subject of Section 3.)

Preface: What are Parameters?

Before diving into the key players and their models, it’s important to understand a fundamental concept that underpins the performance and capabilities of modern AI systems: parameters. These are the core building blocks of neural networks, including Large Language Models (LLMs).

Concept: Parameters

Parameters are the internal values that a model learns during training. They act as “weights” that determine how much importance the model assigns to different patterns in its input data. For example, in a language model, parameters help decide how strongly one word relates to another in a sentence.

The number of parameters in a model is often used as a measure of its size and complexity. Larger models with more parameters tend to perform better on complex tasks because they can capture more nuanced relationships in data. However, recent years have shown that parameter count alone doesn’t tell the whole story – architecture, training data quality, and inference techniques matter just as much.

Beyond Raw Parameter Counts

A Mixture-of-Experts (MoE) model may hold hundreds of billions of total parameters but activate only a small fraction of them per request, so it costs far less to run than its headline size suggests. Meanwhile, a well-trained small model can outperform a much larger one on a specific benchmark. The lesson: parameter count indicates scale, but efficiency, architecture, and training quality determine real-world capability.


The Current Provider Landscape

The AI industry is characterized by intense competition, rapid iteration, and a blurring line between commercial and open-weight offerings. Here are the major players and what they bring to the table.

About the Model Names in This Section

Model examples on this page were verified in August 2026. The AI landscape moves fast. Model names below were verified at the date shown; the concepts they illustrate outlast any particular release. Always check a vendor's current documentation before making a deployment decision.

Specific versions change every few months. What is worth internalizing is each provider’s posture – how they license, where they excel, and what depending on them commits you to.

​

OpenAI

OpenAI logo

OpenAI remains the most recognized name in generative AI. Its current generation folds reasoning into the mainline models as an adjustable effort level rather than shipping a separate reasoning family.

  • GPT-5.6 Sol: The flagship. OpenAI positions it as its workhorse and strongest coding model, aimed at complex reasoning and agentic workflows.
  • GPT-5.6 Terra: The mid-tier option – close to the previous generation's flagship capability at roughly half the cost.
  • GPT-5.6 Luna: The fastest and cheapest of the family, for high-volume and latency-sensitive work.

Strengths: Extensive API ecosystem, robust enterprise support, strong multimodal capability, and adjustable reasoning effort across the whole family.

Access: Via API (api.openai.com), ChatGPT, and Codex, or through Microsoft's Foundry platform (formerly Azure OpenAI Service) for enterprise deployments with enhanced data controls.

Licensing: Proprietary – API access only, no weights released.

Anthropic

Anthropic Claude logo

Anthropic's Claude family is a leading competitor, particularly strong on long-context work, coding, and agentic tasks.

  • Claude Opus 5: The most capable model in the family, with an effort dial that trades cost against depth of reasoning on a per-request basis.
  • Claude Sonnet 5: The balanced option – strong performance at lower cost and higher speed than Opus.
  • Claude Haiku 4.5: Speed-optimized for high-volume, cost-sensitive, and edge-adjacent workloads.

Strengths: Large context windows, a strong safety and interpretability research programme, excellent coding and extended analysis, and Claude Code for agentic development work.

Access: Via API (api.anthropic.com), Amazon Bedrock, or Google Cloud Vertex AI.

Licensing: Proprietary – API access only, no weights released.

Google

Google Gemini logo

Google's Gemini family pairs very large context windows with native multimodality, and Gemma provides an open-weight counterpart.

  • Gemini 3.1 Pro: The generally-available flagship. A one-million-token context window with native multimodality across text, images, video, audio and PDFs.
  • Gemini 3.6 Flash: The current workhorse tier – cheaper and more token-efficient than its predecessor on coding and multi-step tasks.
  • Gemini 3.5 Flash-Lite: The small, low-latency option for high-volume tasks.
  • Gemma: Google's open-weight family for local deployment and fine-tuning.

Strengths: Very large context windows, deep integration with Google Cloud and Workspace, native multimodality including video, and a credible open-weight track through Gemma.

Access: Via Google AI Studio, Vertex AI, or self-hosted using Gemma open weights.

Licensing: Gemini is proprietary; Gemma is released under open weights with use restrictions.

Meta

Meta Llama logo

Meta's position changed sharply in 2026. Its Llama series had been the backbone of the open-source ecosystem; the successor Muse series, from Meta Superintelligence Labs, is proprietary. Meta has said it hopes to open-source future versions.

  • Muse Spark: A natively multimodal reasoning model with tool use, visual chain-of-thought, and multi-agent orchestration. Ships three reasoning modes – Instant, Thinking, and Contemplating (which runs parallel subagents on parts of a complex request).
  • Llama (legacy): Earlier open-weight releases remain available and widely deployed, but are no longer Meta's frontier line.

Strengths: Strong efficiency (Meta reports Llama 4 Maverick-level performance at roughly a tenth of the compute), built-in agentic orchestration, and deployment across Meta's own hardware including smart glasses.

Access: Via Meta's own products and partners. Legacy Llama weights remain self-hostable.

Licensing: Muse Spark is proprietary. Legacy Llama models use the Meta Community License – open weights, but with an acceptable-use policy and a 700-million-monthly-active-user commercial cap. Not OSI-permissive.

A Cautionary Tale About Vendor Strategy

Meta’s Llama series was, for several years, the foundation of the open-weight ecosystem – the default choice for organizations that wanted to self-host. In 2026 Meta moved its frontier work to a proprietary successor.

Existing Llama weights didn’t disappear; anyone running them can keep running them. But teams that had assumed a continuing stream of open frontier models from Meta had to reconsider. When you build on open weights, you are betting on a vendor’s strategy as much as on its technology, and strategies change.

DeepSeek

DeepSeek logo

DeepSeek demonstrated that frontier-adjacent capability does not require the largest budgets, and releases under a genuinely permissive licence.

  • DeepSeek-V4-Pro: The flagship. A Mixture-of-Experts model that activates only a fraction of its parameters per request, making it dramatically cheaper to run than its total parameter count suggests.
  • DeepSeek-V4-Flash: The small, fast variant for cost-sensitive and edge deployment.

Strengths: Exceptional cost efficiency through MoE architecture, a true open-source licence, competitive reasoning via thinking mode, and full self-hosting freedom.

Access: Via API (api.deepseek.com), self-hosted from open weights, or through cloud providers.

Licensing: MIT – genuinely permissive, no usage caps.

What is MoE?

A Mixture-of-Experts model contains many specialized sub-networks but activates only a few for any given token. That means a model with a very large total parameter count can cost far less to run than its size suggests – you pay for the parameters actually used, not the ones sitting idle.

Alibaba (Qwen)

Alibaba Qwen logo

Alibaba's Qwen series combines strong multilingual capability with a wide range of model sizes and a permissive licence.

  • Qwen 3.8-Max: The current flagship of the Qwen line.
  • Qwen 3.x (sizes): Available across a broad parameter range, from models that run on a laptop up to frontier-scale, with strong multilingual support.
  • Qwen-Coder: Specialized coding variants competing with dedicated code models.

Strengths: Industry-leading multilingual support, Apache 2.0 licensing, strong coding and mathematics, and an unusually wide size range.

Access: Via Alibaba Cloud, self-hosted from open weights, or through Hugging Face and cloud providers.

Licensing: Apache 2.0 – genuinely permissive.


Mistral AI

Mistral AI logo

The French company has carved out a niche in efficient, high-quality models with a European data-sovereignty story.

  • Mistral Medium: The balanced tier – strong performance-to-size ratio.
  • Mistral Large: The most capable model in the line, for complex tasks.
  • Mistral Small: Sized to fit on consumer GPUs while staying competitive.

Strengths: High efficiency, excellent performance-to-size ratio, European data sovereignty, and strong multilingual capability.

Access: Via La Plateforme API, self-hosted from open weights, or through cloud providers.

Licensing: Split catalogue – the open models are Apache 2.0, while some commercial models are released under a paid licence. Check per model.

Other Notable Players

Beyond the major providers above, several others play important roles in the ecosystem:

  • Microsoft: Runs the Foundry platform (formerly Azure AI Foundry, and before that Azure OpenAI Service), a single catalogue of thousands of models that spans OpenAI's family alongside open-weight and third-party options, wrapped in enterprise security, compliance and data-residency controls. Also develops the Phi family of small language models in-house.
  • Amazon: Amazon Bedrock provides a unified API across Anthropic, Meta, Mistral and others, with guardrails, fine-tuning and private networking. Amazon also builds its own Nova family, spanning text, image, video and speech-to-speech tiers.
  • Cohere: Enterprise-focused, built around the Command generation models and Rerank retrieval models and its North agent platform. Its differentiator is deployment posture – on-premises, inside your own VPC, or in a sovereign region – for organizations that cannot send data to a public API.
  • xAI: The Grok family has moved from an X platform feature to a frontier-tier competitor with its own API, and is resold through other clouds' model catalogues.

Note how many of these are aggregators rather than model builders. A growing share of enterprise AI consumption reaches a model through someone else’s catalogue – which is convenient, but means your security review has to cover the platform as well as the model.


The Open-Source vs. Closed-Source Shift

One of the most significant developments in recent years is the narrowing gap between open-weight and proprietary models. This shift has profound implications for how organizations approach AI deployment.

Deep Dive: The Open-Source Revolution

What changed:

  • Open-weight models have repeatedly demonstrated reasoning quality matching proprietary ones
  • Leading open-weight models compete with frontier proprietary models on many benchmarks
  • Qwen, Mistral, and other open-weight providers offer commercially licensed models
  • The cost of fine-tuning and deploying open models has dropped dramatically

Why it matters:

  • Organizations can now choose based on data privacy, cost, and control rather than capability alone
  • Security teams can inspect and test model weights directly (impossible with closed APIs), though weights alone don’t reveal training data
  • Fine-tuning for domain-specific tasks is accessible to any team with modest GPU resources
  • No vendor lock-in – switch between providers or run multiple models

The trade-off:

  • Proprietary frontier models still often lead on the most complex tasks
  • Closed providers handle infrastructure, scaling, and updates
  • Open models require operational expertise for deployment and maintenance

Lifecycle of an LLM

Training Process Overview

Training an LLM is like teaching a student by exposing them to an enormous library of books. The model learns patterns, relationships, and structures in language by processing massive datasets. This process involves several steps:

  1. Data Collection and Preprocessing: Text data is gathered from diverse sources such as books, websites, and articles. The data is cleaned, tokenized (split into smaller units), and converted into numerical representations.
  2. Model Configuration: Parameters like the number of layers, attention heads, and learning rates are set. These define the model’s architecture and training dynamics.
  3. Optimization: Using algorithms like gradient descent, the model adjusts its parameters to minimize errors in predicting token sequences.
Concept: Training

Training is the process of teaching an LLM by exposing it to vast amounts of text data. The model learns to predict the next token in a sequence based on context, gradually improving its understanding of language.


Fine-Tuning Methodology (Optional)

Once trained on general data, an LLM can be fine-tuned for specific tasks or domains. For example:

  • Instruction Fine-Tuning: Teaching the model how to respond to specific prompts (e.g., summarization or question-answering).
  • Parameter-Efficient Fine-Tuning (PEFT): Updating only a small subset of parameters (e.g., using techniques like LoRA) to adapt the model without retraining it entirely.
  • RLHF (Reinforcement Learning from Human Feedback): Training the model to align its outputs with human preferences – this is how models like ChatGPT learn to be helpful and refuse harmful requests.

Fine-tuning allows organizations to customize models for applications like legal document analysis or customer support while maintaining efficiency.

Concept: Fine-tuning

Fine-tuning involves adapting a pre-trained LLM to perform specific tasks or operate within particular domains by retraining on targeted datasets.


Inference

Inference is where all the training pays off – it’s when a trained LLM generates outputs based on new inputs. During inference:

  • The input text is tokenized and passed through the model.
  • The model predicts the most likely next tokens based on its learned patterns.
  • These tokens are decoded back into human-readable text.

Optimizing inference for speed and efficiency is critical for real-time applications like chatbots or virtual assistants.

Concept: Inference

Once trained, the model applies its knowledge to new inputs – like answering questions or generating text. This phase prioritizes speed and efficiency. When you generate text with an LLM – for instance while using ChatGPT – you are using the model in inference mode.


Types of AI Models: Foundation and Specialized

AI models can be broadly categorized based on their purpose and training process. Understanding these distinctions is crucial to grasp how modern AI systems are built and deployed.

Foundation Models

Foundation models are large-scale, pre-trained systems designed to handle a wide range of tasks. They are trained on massive, diverse datasets using self-supervised learning, enabling them to generalize across domains. These models serve as a starting point for further customization or direct application.

  • Key Characteristics:

    • General-purpose and adaptable.
    • Pre-trained on diverse datasets spanning multiple domains.
    • Can perform many tasks out-of-the-box (e.g., text generation, image analysis).
  • Examples: The flagship models from each major provider are foundation models – see the provider landscape above for current names. Whether proprietary (API-only) or open-weight (self-hostable), they share the same defining trait: broad pre-training that makes them useful across many tasks without further modification.

Foundation models are often fine-tuned or adapted for specific applications, which leads us to the next type.

Specialized Models: Fine-Tuned vs. Custom-built

Specialized models are AI systems designed to excel at specific tasks or domains. They can be developed in two primary ways: by fine-tuning a foundation model or by building a model entirely from scratch. Both approaches have distinct advantages, trade-offs, and use cases.

Fine-Tuned Models

These are derived from pre-trained foundation models by further training them on smaller, task-specific datasets. Fine-tuning leverages transfer learning, allowing the model to retain general knowledge from its pretraining while adapting to the nuances of a particular domain or task. This approach is cost-effective and efficient, as it requires significantly fewer resources than training a model from scratch.

Key Characteristics:

  • Built on top of foundation models.
  • Require smaller, domain-specific datasets.
  • Cost-effective and efficient compared to scratch-built models.
  • Moderately adaptable for related tasks.

Examples:

  • MedGemma: Fine-tuned from Google’s open-weight Gemma family on de-identified medical text and imaging – chest X-rays, dermatology, histopathology – for clinical research and health AI development.
  • Qwen-Coder: Adapted from Alibaba’s general-purpose Qwen models for programming tasks like code generation.
  • Security-domain models: Foundation models fine-tuned on threat intelligence, malware samples, and vulnerability data, used for alert triage and incident summarization. Most major security vendors now maintain one.

Fine-tuning is ideal when a foundation model provides sufficient baseline capabilities but needs customization for domain-specific applications.

Custom-built Models

These are developed entirely from scratch, tailored to a specific problem or domain. Custom-built models do not rely on pre-trained weights, making them ideal for scenarios where proprietary data, extreme precision, or unique architectures are required. However, this approach is resource-intensive and demands substantial expertise.

Key Characteristics:

  • Designed specifically for one task or domain.
  • Not general-purpose; lacks adaptability beyond its training focus.
  • Requires significant computational resources and expertise.
  • Offers unmatched precision for niche applications.

Examples:

  • DeepMind AlphaFold: Built to predict the structure of proteins and their interactions with other molecules, on an architecture that shares nothing with a language model. It reshaped structural biology.
  • BloombergGPT: A 50B-parameter financial language model trained from scratch on Bloomberg’s proprietary archive alongside general text, for sentiment analysis, entity recognition and question answering over financial documents.
  • DeepMind GenCast: A weather model trained from scratch on decades of atmospheric data. It produces a 15-day global forecast in minutes on a single accelerator, beating the leading physics-based ensemble system that needs hours on a supercomputer.

Custom-built models are indispensable when extreme precision or unique domain requirements cannot be met by existing foundation models.


Concept: Horizontal vs. Vertical Specialization

The distinction between horizontal and vertical specialization applies to both fine-tuned and scratch-built models:

  • Horizontal Models: Designed for broad, cross-domain tasks. Most frontier foundation models are horizontal, but scratch-built horizontal systems exist too – national and regional labs regularly train general-purpose models from scratch for sovereignty or language-coverage reasons rather than to specialize them.
  • Vertical Models: Tailored to specific industries or tasks. These can be fine-tuned versions of foundation models (e.g., MedGemma) or scratch-built systems (e.g., BloombergGPT).

Horizontal models prioritize versatility across domains, while vertical models emphasize precision and domain expertise.


Key Differences Between Model Types

Feature Foundation Models Fine-Tuned Models Custom-built Models
Scope General-purpose Domain-specific Task-specific
Training Data Diverse datasets Domain-specific datasets Focused task/domain-specific data
Flexibility Highly adaptable Moderately adaptable Not adaptable
Development Cost High (pre-training stage) Low (retraining only) Very high (built from scratch)
Examples Frontier models from major providers MedGemma, security-domain fine-tunes AlphaFold, GenCast, BloombergGPT

Why These Distinctions Matter

Understanding these distinctions is crucial for evaluating AI solutions:

  • Fine-tuned models offer a balance of adaptability and efficiency by leveraging pre-trained knowledge.
  • Scratch-built models provide unmatched precision when domain requirements exceed what foundation models can deliver.
  • Horizontal vs. vertical specialization helps clarify whether a model is designed for broad applicability or tailored to a specific industry or task.

By recognizing these differences, organizations can make informed decisions about which type of model best suits their needs – whether it’s developing a general-purpose tool or solving a highly specific problem.


The Size Axis

Foundation, fine-tuned and custom-built describe what a model is for. Size is a separate axis, and it cuts across all three: providers ship their families in tiers, and where a model sits on that spectrum decides as much about your architecture as its capability does.

Tier Typical scale Where it runs What it’s for
Small Language Models (SLMs) ~1B-15B parameters Phones, laptops, edge and embedded devices – offline if needed Focused, well-defined tasks: classification, extraction, translation, on-device assistants
Mid-tier Roughly 15B-100B parameters A single well-specified server, or a hosted API The general workhorse – most production traffic, where cost per request matters
Frontier Very large, often MoE GPU clusters, or an API to someone else’s cluster The hardest reasoning, long-context analysis, and agentic work

Every provider tab above shows this pattern: a flagship, a balanced middle, and a fast small option. That is not a coincidence – it exists because the right answer changes with the workload.

The Size Tier Is a Deployment Decision

The SLM tier is the one that most often changes an architecture rather than just a line item. A model small enough to run on the device is a model whose inputs never leave the device – which resolves connectivity, latency and data-residency constraints in one move, at the cost of capability on open-ended work.

So the useful question when selecting a model is rarely “which one is best?” but “what is the smallest one that reliably does this job?” Answering it well is usually worth more than any capability difference between the frontier options.

Microsoft's Phi-4-mini (3.8B), Google's Gemini 3.5 Flash-Lite, Anthropic's Claude Haiku 4.5, and DeepSeek-V4-Flash.

SLMs are introduced in Section 1; Section 3 works through on-device and self-hosted deployment in detail, including the security consequences of each.


Extended Reasoning

Extended reasoning is one of the most significant developments of recent years – and a useful lesson in how quickly a product category can dissolve. It arrived as a separate class of model: vendors shipped dedicated “reasoning models” alongside their standard ones, and you chose between them. That distinction has largely gone. Reasoning is now an adjustable effort or thinking setting on mainline models, so the choice is no longer which model but how hard it should think about this request.

The underlying mechanism is unchanged either way. Rather than producing an answer in a single pass, the model first generates an extended chain of intermediate tokens – breaking the problem into steps, evaluating approaches, and sometimes backtracking when it detects an error in its own working – and only then writes the visible response. This is test-time compute: spending more effort at the moment of answering, rather than baking more capability in during training.

A Dial, Not a Species

Because this is a setting rather than a product tier, the comparison that matters is between responses, not between models. The same model answers a trivial question cheaply and a hard one expensively, depending on where you set the dial:

OpenAI's GPT-5.6 Sol at high reasoning effort, Anthropic's Claude Opus 5 with its effort dial, Google's Gemini 3.x thinking levels, and DeepSeek-V4 in thinking mode.

Open-weight models offer this too, so extended reasoning is not exclusive to proprietary APIs.


Standard vs. Extended Reasoning

Aspect Standard response Extended reasoning
What happens The answer is generated directly, in a single pass. The model works through intermediate steps privately, then answers.
Best use cases Summarization, content creation, classification, basic Q&A. Multi-step reasoning: debugging, mathematics, research synthesis, planning.
Strengths Fast, cheap, and entirely sufficient for most traffic. Markedly better on problems where a wrong intermediate step ruins the answer.
Limitations Weaker where a problem genuinely needs to be decomposed. Slower, consumes many more tokens, and can overthink simple tasks.

The practical consequence is a cost decision, not a capability one. Reasoning tokens are billed like any other, so turning the dial up on every request is an easy way to multiply a bill for no benefit. Turn it up where a wrong intermediate step would ruin the result; leave it down everywhere else.


Real-World Applications

Application Dial down Dial up
Content Generation Writing marketing copy Generating a detailed technical report.
Customer Support Answering FAQs Resolving multi-step troubleshooting queries.
Data Analysis Extracting key statistics Analyzing trends across multiple datasets.
Education Flashcard-style Q&A Teaching complex concepts step-by-step.
Software Engineering Code completion and suggestions Debugging complex systems and architectural planning.
Key Takeaways
  • The provider landscape is intensely competitive and genuinely international, spanning the US, China and Europe, with commercial and open-weight offerings blurring into each other
  • Models are categorized by purpose – foundation, fine-tuned, custom-built – and separately by size, from on-device SLMs up to frontier systems
  • The LLM lifecycle consists of training (learning patterns from data), optional fine-tuning (domain adaptation), and inference (generating outputs)
  • Extended reasoning is a dial on a mainline model, not a separate class of model: a cost-versus-depth decision made per request
  • The open-source vs. closed-source gap has narrowed dramatically, enabling organizations to choose based on data privacy and control rather than capability alone – but open weights are a bet on a vendor’s strategy, and strategies change

Test Your Knowledge

Ready to test your understanding of key AI players and models? Head to the quiz to check your knowledge.


Up next

Understanding the key players and their offerings is just the first step. In the next section, we’ll dive deeper into deployment considerations for these models, including security implications, model selection criteria, and different deployment options. We’ll explore what it takes to implement LLMs in real-world enterprise scenarios securely and effectively.