This video introduces weight quantization and demonstrates how to reduce the size of large language models using 8-bit quantization. The notebook includes code for absmax and zeropoint quantization, as well as generating text and calculating perplexity for both the original and quantized models. The notebook also includes code for loading a pre-trained model in 8-bit format and comparing the dequantized weights to the original weights.
This is an autogenerated video based on Jupyter Notebooks. Do you want to generate your own videos? Go to and try our tool for free: https://jupyvideo.feltlabs.ai
This tutorial is based on Jupyter notebook: https://github.com/mlabonne/llm-cours...
In questa pagina del sito puoi guardare il video online LLM's Weight Quantization Explained della durata di ore minuti seconda in buona qualità , che l'utente ha caricato FELT Labs 01 gennaio 1970, condividi il link con amici e conoscenti, su youtube questo video è già stato visto 217 volte e gli è piaciuto 3 spettatori. Buona visione!