コンテンツへスキップ
用語集に戻る
アーキテクチャ

Vision Transformer (ViT)

定義

An adaptation of the transformer architecture to image data, treating fixed-size image patches as tokens. ViTs now outperform convolutional networks on many computer vision benchmarks and are used in medical imaging, satellite analysis, and industrial quality control.

関連サービス

AIの理解にお困りですか?

Physical AI 適合性コールをご予約いただき、これらのAI概念が貴社の業界や課題にどう適用されるかをご相談ください。

Vision Transformer (ViT) | AI用語集 — LLM、RAG、ファインチューニング&主要概念