Skip to content
Back to Glossary
Architecture

Vision Transformer (ViT)

Definition

An adaptation of the transformer architecture to image data, treating fixed-size image patches as tokens. ViTs now outperform convolutional networks on many computer vision benchmarks and are used in medical imaging, satellite analysis, and industrial quality control.

Knowing the Terms Is Step One. Applying Them Is Step Two.

Book a Physical AI Fit Call to discuss how these AI concepts translate to your specific industry and business challenges.

Vision Transformer (ViT) | AI Glossary — LLM, RAG & 20+ Key Terms Explained