An adaptation of the transformer architecture to image data, treating fixed-size image patches as tokens. ViTs now outperform convolutional networks on many computer vision benchmarks and are used in medical imaging, satellite analysis, and industrial quality control.
احجز مكالمة تقييم ملاءمة Physical AI لمناقشة كيفية تطبيق مفاهيم الذكاء الاصطناعي هذه على قطاعك وتحدياتك.