Your first local LLM
Llama 3.1 8B has a measured Q4_K_M download of 4.92 GB and needs about 6.4 GB of memory at 4k context. Find a fit for your hardware below, then follow the setup steps.
Find a local model, then run it
Start with your device and job. Recommendations use estimated Q4 memory and measured download size. They do not rank model quality or guarantee speed.
Choose your own hardware below. The initial device is an example, not detected hardware.
Everyday questions, drafting, and summaries. Context is how much chat and text the runtime keeps available.
Popular models that fit
Ordered by recorded Ollama pulls, not quality scores.
Set up Llama 3.1 8B
- Install and open LM Studio.
- Search for bartowski/Meta-Llama-3.1-8B-Instruct-GGUF, then choose the Q4_K_M GGUF file.
- In My Models, open the gear icon and set the context window to 4000 tokens before loading it.
- Open a chat and test one task you actually use.
LM Studio download guide · Per-model settings · LM Studio chat guide · Check runtime system requirements · Other local tools
Inspect this device estimateDid it work?
Your voluntary response records one aggregate page view. It does not include your device or model choice.
Memory fit is an estimate, not a tested runtime result. Read the calculation method or compare recorded benchmarks.