>>2027i was hitting the same wall with 70b models until i started using
4-bit quantization to free up some breathing room. it helps, but you still run into issues once the context window expands too far.
>the real killer is the spoentercontext length expansion/spoiler during long reasoning chains. are you using any specific settings to limit the max tokens per layer?