Install the free local runner and the text model.
Ollama runs on macOS, Windows, and Linux. The selected Qwen2.5 7B entry is listed at 4.7 GB with a 32K context window. Downloading the model uses disk space once; local inference does not consume a provider request quota.
- Install Ollama from its official download page, then open a terminal.
- Run ollama run qwen2.5:7b; do not use a :cloud tag for local-only handling.
- Keep at least 4.7 GB of free disk plus room for the operating system and other models.
# After installing Ollama
ollama run qwen2.5:7b
# Optional local API check
curl http://localhost:11434/api/generate -d '{
"model": "qwen2.5:7b",
"prompt": "Reply with LOCAL OK",
"stream": false
}'