58Fit score
Small open-weight model for lightweight deployment.
Model details →Small models (≤ 8B params) that run on consumer GPUs or even phones. Trade-off: lower raw quality, but private, offline, and instant.
Small open-weight model for lightweight deployment.
Model details →Low-cost model for high-volume API workloads.
Model details →Fast reasoning model optimized for cost/performance.
Model details →| # | Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|---|
| 1 | Llama 3.1 8B Instruct | Meta | 128K Long context (128K+) | Custom | Variable |
| 2 | GPT-4o mini | OpenAI | 128K Long context (128K+) | Custom | Variable |
| 3 | o3-mini | OpenAI | 200K Long context (128K+) | Custom | Variable |
| 4 | Jamba 1.5 Mini | AI21 | 256K Long context (128K+) | Custom | Variable |
| 5 | Mistral Small | Mistral AI | 128K Long context (128K+) | Custom | Variable |
Most 1–8B models run on a recent MacBook with 16GB+ RAM via Ollama or llama.cpp. Quality is well below frontier models but private and offline.