The memory is plausible, but the INT8 / FP8 + offload profile needs an OpenMayhem benchmark first.
- Model size
- 8.9B · 40-step generation
- Model format
- INT8 / FP8 + offload
- Runs with
- Diffusers
This profile is provisional and needs a local OpenMayhem benchmark before serving.