Skip to content

Your first local LLM

Llama 3.1 8B has a measured Q4_K_M download of 4.92 GB and needs about 6.4 GB of memory at 4k context. Find a fit for your hardware below, then follow the setup steps.

Guided setupMemory estimate

Find a local model, then run it

Start with your device and job. Recommendations use estimated Q4 memory and measured download size. They do not rank model quality or guarantee speed.

Choose your own hardware below. The initial device is an example, not detected hardware.

Everyday questions, drafting, and summaries. Context is how much chat and text the runtime keeps available.

Popular models that fit

Ordered by recorded Ollama pulls, not quality scores.

Set up Llama 3.1 8B

  1. Install and open LM Studio.
  2. Search for bartowski/Meta-Llama-3.1-8B-Instruct-GGUF, then choose the Q4_K_M GGUF file.
  3. In My Models, open the gear icon and set the context window to 4000 tokens before loading it.
  4. Open a chat and test one task you actually use.

LM Studio download guide · Per-model settings · LM Studio chat guide · Check runtime system requirements · Other local tools

Inspect this device estimate

Did it work?

Your voluntary response records one aggregate page view. It does not include your device or model choice.

Memory fit is an estimate, not a tested runtime result. Read the calculation method or compare recorded benchmarks.