Monthly cost at list price for five common workloads. Token counts per unit are stated so you can scale them to your own traffic. A workload that does not fit a model's context window is marked rather than priced.
Workload (per month)
Google: Gemini 3.1 Flash Lite
MiniMax: MiniMax M3 (free)
Difference
Customer-support chatbot
50,000 conversations — each a 6-turn conversation (1,500 in / 400 out tokens)
$48.75
$0
MiniMax: MiniMax M3 (free) saves $48.75
RAG search over documents
20,000 questions — each a question answered from ~8 retrieved passages (6,000 in / 500 out tokens)
$45.00
$0
MiniMax: MiniMax M3 (free) saves $45.00
Coding / tool-using agent
2,000 tasks — each a multi-step task with tool calls and file context (60,000 in / 6,000 out tokens)
$48.00
$0
MiniMax: MiniMax M3 (free) saves $48.00
Batch summarisation
10,000 documents — each a 6,000-word report summarised to a page (8,000 in / 600 out tokens)
$29.00
$0
MiniMax: MiniMax M3 (free) saves $29.00
Long-document analysis
500 documents — each a 150-page contract or codebase read in one pass (100,000 in / 2,000 out tokens)
$14.00
$0
MiniMax: MiniMax M3 (free) saves $14.00
Capabilities side by side
Capability
Google: Gemini 3.1 Flash Lite
MiniMax: MiniMax M3 (free)
What it lets you do
Tool / function calling
✓ Yes
✓ Yes
call your APIs and run agent loops
Structured output (JSON schema)
✓ Yes
✓ Yes
return JSON that validates against your schema
Reasoning / thinking mode
✓ Yes
✓ Yes
spend extra tokens thinking before answering hard problems
Image input
✓ Yes
✓ Yes
read screenshots, charts and scanned pages
Audio input
✓ Yes
—
take speech or audio directly
File / PDF input
✓ Yes
—
accept documents without your own parsing step
Image output
—
—
generate images, not just text
Built-in web search
✓ Yes
—
answer from live web results
Seed (reproducible sampling)
✓ Yes
✓ Yes
repeat a generation for tests and evals
Log-probabilities
—
—
read token confidence for classification and scoring
From each model's published parameters and input/output types. “—” means not advertised, not proven absent.
Price history
Google: Gemini 3.1 Flash Lite: price unchanged since we started tracking it on 2026-08-06.
MiniMax: MiniMax M3 (free): price unchanged since we started tracking it on 2026-08-26.
Which should you choose?
Choose Google: Gemini 3.1 Flash Lite if…
you need audio input — to take speech or audio directly; MiniMax: MiniMax M3 (free) does not advertise it
you need file / PDF input — to accept documents without your own parsing step; MiniMax: MiniMax M3 (free) does not advertise it
you need built-in web search — to answer from live web results; MiniMax: MiniMax M3 (free) does not advertise it
Choose MiniMax: MiniMax M3 (free) if…
cost matters at scale — it is 100% cheaper on a typical 3:1 input/output mix ($0 vs $48.75 a month for 50,000 support chats)
you generate long outputs in one go — up to 944K tokens (about 707,800 words) per response
you want to prototype at zero cost — a free variant is listed
Frequently asked questions
Is Google: Gemini 3.1 Flash Lite better than MiniMax: MiniMax M3 (free)?
Neither wins outright: Google: Gemini 3.1 Flash Lite is the pick when you need audio input, MiniMax: MiniMax M3 (free) when cost matters at scale. In short: Google: Gemini 3.1 Flash Lite when you need audio input, you need file / PDF input, you need built-in web search; MiniMax: MiniMax M3 (free) when cost matters at scale, you generate long outputs in one go, you want to prototype at zero cost.
Which is cheaper, Google: Gemini 3.1 Flash Lite or MiniMax: MiniMax M3 (free)?
MiniMax: MiniMax M3 (free) is cheaper: $0 vs $0.56 per 1M tokens at a typical 3:1 input/output mix, 100% less. For 50,000 support conversations a month that is $0 against $48.75.
How much do Google: Gemini 3.1 Flash Lite and MiniMax: MiniMax M3 (free) cost per million tokens?
Google: Gemini 3.1 Flash Lite costs $0.25 per 1M input tokens and $1.50 per 1M output tokens; MiniMax: MiniMax M3 (free) costs $0 input and $0 output. Output tokens are billed separately, so long answers shift the gap.
Which has the bigger context window, Google: Gemini 3.1 Flash Lite or MiniMax: MiniMax M3 (free)?
Both take 1.0M tokens per request — roughly 1,573 pages of text.
Do Google: Gemini 3.1 Flash Lite and MiniMax: MiniMax M3 (free) support function calling and JSON output?
Can Google: Gemini 3.1 Flash Lite or MiniMax: MiniMax M3 (free) read images?
Yes, both accept image input.
Is Google: Gemini 3.1 Flash Lite or MiniMax: MiniMax M3 (free) free to use?
MiniMax: MiniMax M3 (free) has a free variant (minimax/minimax-m3:free). Free variants usually carry tighter rate limits.
What is the maximum output length of Google: Gemini 3.1 Flash Lite and MiniMax: MiniMax M3 (free)?
Google: Gemini 3.1 Flash Lite can write up to 66K tokens (about 49,200 words) per response; MiniMax: MiniMax M3 (free) up to 944K (about 707,800 words).
When were Google: Gemini 3.1 Flash Lite and MiniMax: MiniMax M3 (free) released?
Google: Gemini 3.1 Flash Lite was listed on 2026-05-07; MiniMax: MiniMax M3 (free) on 2026-05-31.
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...