Comfortable on the SHQ8 MTP GGUF profile.
Edition: Quantized edition by Wepiqx
- Model size
- 9B · up to 1M context
- Model format
- SHQ8 MTP GGUF
- Runs with
- llama.cpp
This profile includes practical headroom, but context length and concurrency still affect memory use.