Expected to run with a constrained Q4 GGUF profile.
Edition: Uncensored edition by HauHau
- Model size
- 35B / 3B active · 262K max
- Model format
- Q4 GGUF
- Runs with
- llama.cpp
This profile includes practical headroom, but context length and concurrency still affect memory use.