Techniques

Mixture of Experts (MoE)

Definition

A model architecture where different sub-networks ("experts") specialise in different types of inputs, and a gating network routes each token to the most relevant experts. MoE enables very large model capacity at lower inference cost—Mixtral and GPT-4 are believed to use this approach.

Related Terms

Transformer

The dominant neural network architecture for language, vision, and multimodal AI, introduced in the 2017 "Attention Is All You Need" paper. Transformers use self-attention to process all tokens in parallel, enabling training on internet-scale data and powering every major LLM in use today.

Sparse Model

A model architecture that activates only a subset of its parameters for any given input, rather than the full network. Sparse models—enabled by Mixture of Experts designs—achieve larger total capacity while keeping per-inference compute manageable.

Inference

The process of running a trained model on new data to produce predictions or generated outputs. Inference cost and latency are the dominant operational concerns in production AI, particularly for large generative models that can cost cents per request at scale.

Knowing the Terms Is Step One. Applying Them Is Step Two.

Book a discovery call to discuss how these AI concepts translate to your specific industry and business challenges.