Pytorch for Beginners #28 | Transformer Model: Multiheaded Attention - Optimize Basic Implementation

Pubblicato il: 21 novembre 2021
sul canale di: Makeesy AI
1,564
31

Transformer Model: Multiheaded Attention - Optimize Basic Implementation

In this tutorial, we’ll optimize the basic implementation of multiheaded attention. We’ll discuss in detail that why using separate query, key, and value matrices for various head is inefficient both coding-wise as well as computation-wise. We’ll optimize the implementation by packing the weights of query, key, and value matrices for various heads. Specifically, we’ll use tensor reshaping to achieve the required behavior. For more details on view(), and contiguous() methods, watch the related tutorials, links given below.

Tensor reshaping –    • Pytorch for Beginners: #6 | Modify Te...  

Contiguous vs non-contiguous tensors -    • Pytorch for Beginners: #7 | Contiguou...  

The code used in this tutorial is available here- https://github.com/makeesyai/makeesy-...

Chapters-
0:00 - Introduction
0:20 - Limitations of basic implementation
1:15 - Packing Query, Key and Value weights for various heads
5:50 - Reshaping Query, Key, and Value segregating various heads
9:38 - Unify heads reshaping weighted values
13:53 - Compare output with basic implementation
16:05 - Next

#pytorch #toturial #multiheaded #attention #optimized #implementation


In questa pagina del sito puoi guardare il video online Pytorch for Beginners #28 | Transformer Model: Multiheaded Attention - Optimize Basic Implementation della durata di ore minuti seconda in buona qualità , che l'utente ha caricato Makeesy AI 21 novembre 2021, condividi il link con amici e conoscenti, su youtube questo video è già stato visto 1,564 volte e gli è piaciuto 31 spettatori. Buona visione!