Expected to run with a constrained Q4_K GGUF · reduced context profile.
Edition: Abliterated edition by InternScience Agents-A1
- Model size
- 35B / 3B active · 262K max
- Model format
- Q4_K GGUF · reduced context
- Runs with
- llama.cpp
This profile includes practical headroom, but context length and concurrency still affect memory use.