Tutorial 15: Vision Transformers

Published: 09 October 2021
on channel: UvA Deep Learning course
4,983
50

In this tutorial, we will take a closer look at a recent new trend: Transformers for Computer Vision. Since Alexey Dosovitskiy et al. (https://openreview.net/pdf?id=YicbFdNTTy) successfully applied a Transformer on a variety of image recognition benchmarks, there have been an incredible amount of follow-up works showing that CNNs might not be optimal architecture for Computer Vision anymore. But how do Vision Transformers work exactly, and what benefits and drawbacks do they offer in contrast to CNNs? We will answer these questions by implementing a Vision Transformer ourselves, and train it on the popular, small dataset CIFAR10. We will compare these results to popular convolutional architectures such as Inception, ResNet and DenseNet. This notebook is part of a lecture series on Deep Learning at the University of Amsterdam. The full list of tutorials can be found at https://uvadlc-notebooks.rtfd.io.
Link to the notebook: https://uvadlc-notebooks.readthedocs....

00:00 Introduction
02:50 Transformers for Vision
06:40 Vision Transformer Architecture
11:20 Experiments


On this page of the site you can watch the video online Tutorial 15: Vision Transformers with a duration of hours minute second in good quality, which was uploaded by the user UvA Deep Learning course 09 October 2021, share the link with friends and acquaintances, this video has already been watched 4,983 times on youtube and it was liked by 50 viewers. Enjoy your viewing!