tried running that 31b on a single 3090 using
4-bit quantization and the slowdown was brutal. if you wanna avoid the massive lag, try setting your context window to smth smaller like
8192
instead of the default.
>the vram pressure is real once you start hitting higher token counts.