quantization

Just a note: After quantization, AI parameters can typically be compressed to 1/10th of their size.

Take the Llama 2, which we mentioned before, requiring roughly 140 GB RAM. Post-quantization, it could potentially run on a 16GB RAM computer.

Meaning, there’s no need for the pricier Mac Pro. 🖥️

#post_ai_geneneration