Comfortable on the Q4 GGUF + projector profile.
Edition: Uncensored edition by HauHau
- Model size
- 9B dense · 262K max
- Model format
- Q4 GGUF + projector
- Runs with
- llama.cpp
This profile includes practical headroom, but context length and concurrency still affect memory use.