← All demos

Running a local model: the whole pipeline

Four pieces. llama.cpp, Ollama, and LM Studio are all just different front doors to this exact flow.

1 Β· Disk

Download the model

πŸ“„ qwen3.6-35b-a3b.gguf 20.4 GB
2 Β· Memory

Load it into RAM

0 GB of 48 GB fast memory
3 Β· Server

Start the inference server

localhost:1234
not running
4 Β· Agent

Point Little Coder at it

$ little-coder \
  --model lmstudio/local-model

Use ← β†’ to step. Numbers shown: Qwen 3.6 35B-A3B at 4-bit with a 32k context on a 64 GB MacBook (M3 Max β€” ~48 GB usable by the GPU).