Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction follo…
Model details →Best multilingual AI models
Includes Chinese-strong models (Qwen, DeepSeek, GLM, MiniMax, Kimi, ERNIE), European leaders (Mistral, Cohere), and the global frontier APIs.
78Fit score
78Fit score
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose …
Model details →78Fit score
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It.…
Model details →Top 25 models for non-English languages
| # | Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|---|
| 1 | Qwen: Qwen3 30B A3B Instruct 2507 | Qwen | 262K Long context (128K+) | $0.05 | Budget |
| 2 | Qwen: Qwen3 235B A22B Instruct 2507 | Qwen | 262K Long context (128K+) | $0.09 | Budget |
| 3 | Qwen: Qwen3 Max | Qwen | 262K Long context (128K+) | $0.78 | Budget |
| 4 | Qwen: Qwen3 Next 80B A3B Instruct | Qwen | 262K Long context (128K+) | $0.09 | Budget |
| 5 | Qwen: Qwen3 Next 80B A3B Instruct (free) | Qwen | 262K Long context (128K+) | $0.00 | Budget |
| 6 | Qwen: Qwen3 30B A3B | Qwen | 131K Long context (128K+) | $0.12 | Budget |
| 7 | Qwen: Qwen2.5-VL 7B Instruct | Qwen | 33K Short/standard context | $0.20 | Budget |
| 8 | Cohere: Command A | Cohere | 256K Long context (128K+) | $2.50 | Standard |
| 9 | Xiaomi: MiMo-V2-Flash | Xiaomi | 262K Long context (128K+) | $0.09 | Budget |
| 10 | Cohere: Command R (08-2024) | Cohere | 128K Long context (128K+) | $0.15 | Budget |
| 11 | meta-llama/Llama-3.2-3B-Instruct | Meta | 131K Long context (128K+) | $0.05 | Budget |
| 12 | Meta: Llama 3.2 3B Instruct (free) | Meta | 131K Long context (128K+) | $0.00 | Budget |
| 13 | Meta: Llama 3.3 70B Instruct | Meta | 131K Long context (128K+) | $0.10 | Budget |
| 14 | Mistral: Mistral Nemo | Mistral AI | 131K Long context (128K+) | $0.02 | Budget |
| 15 | Mistral: Mistral Small 3.1 24B (free) | Mistral AI | 128K Long context (128K+) | $0.00 | Budget |
| 16 | meta-llama/Llama-3.2-1B-Instruct | Meta | 60K Short/standard context | $0.03 | Budget |
| 17 | Meta: Llama 3.3 70B Instruct (free) | Meta | 66K Short/standard context | $0.00 | Budget |
| 18 | DeepSeek Reasoner (legacy id) | DeepSeek | 1M Ultra context (1M+) | Custom | Variable |
| 19 | DeepSeek: DeepSeek V4 Flash 0423 | DeepSeek | 1.0M Ultra context (1M+) | $0.14 | Budget |
| 20 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 21 | DeepSeek: DeepSeek V4 Pro | DeepSeek | 1.0M Ultra context (1M+) | $0.43 | Budget |
| 22 | Google: Gemma 3n 2B (free) | 8K Short/standard context | $0.00 | Budget | |
| 23 | BAAI/bge-reranker-v2-m3 | Hugging Face | Unknown Unknown context | Custom | Variable |
| 24 | baidu/Unlimited-OCR | Hugging Face | Unknown Unknown context | Custom | Variable |
| 25 | cross-encoder/mmarco-mMiniLMv2-L12-H384-v1 | Hugging Face | Unknown Unknown context | Custom | Variable |
Frequently asked questions
Which model is best for Chinese?
Qwen, GLM (Z.AI), Kimi (Moonshot), DeepSeek, and ERNIE (Baidu) all train heavily on Chinese corpora. For mixed Chinese/English, Qwen and GLM are strong defaults.
Related use cases
AI models for codingAI reasoning modelsAI vision modelslong-context AI models (128K+)AI models for function callingAI models for structured / JSON outputsAI models for agentsfree AI modelsCheapest AI models per tokenembedding modelssmall / on-device AI modelsAI image generation modelsAI voice / audio modelsAI models for mathAI models for writingopen-source AI modelsAI models for RAGAI models for summarizationAI models for translationAI models for enterprise