audio model · orpheus · macOS
Can I run Orpheus 3B on Apple M4 (16GB)?
Yes. Orpheus 3B runs on Apple M4 (16GB) at Q4_K_M GGUF (~4 GB of ~10.5 GB usable).
Runs at Q4_K_M GGUF using ~4 GB of ~10.5 GB usable.
- Peak memory
- ~4 GB
- Usable on device
- ~10.5 GB
- Device memory
- 16 GB
- Quant
- Q4_K_M GGUF
How to run it
Use llama.cpp or LM Studio at Q4_K_M GGUF. It is light enough to run on CPU; a GPU just makes it faster.
- Type
- Text to speech
- Parameters
- 3B
- Peak memory
- ~4 GB at Q4_K_M GGUF
- License
- Apache-2.0
- Memory
- 16 GB unified
- Usable for weights
- ~10.5 GB
- Power draw
- ~65 W
- Best runtime
- Ollama (MLX backend, preview) / MLX direct
You could also run
Run Orpheus 3B on other hardware
FAQ
Can Apple M4 (16GB) run Orpheus 3B?
Yes. Orpheus 3B runs on Apple M4 (16GB) at Q4_K_M GGUF (~4 GB of ~10.5 GB usable).
How much memory does Orpheus 3B need?
Apple M4 (16GB) has room to spare. At Q4_K_M GGUF the realistic peak is ~4 GB of memory.
What do I use to run Orpheus 3B locally?
Orpheus 3B runs in llama.cpp or LM Studio (among others). It runs on CPU, so no GPU is required.
Sources
-
apple.com · 1 source
-
developer.apple.com · 1 source
-
github.com · 3 sources
-
huggingface.co · 2 sources
-
lmstudio.ai · 1 source
-
support.apple.com · 4 sources
VRAM figures are sourced peak-usage anchors at the noted quant; catalog updated 2026-10-05. See methodology.