MiniMax rolled out a new text model, M3.1-Flash-Preview, inside its coding product MiniMax Code on September 27. The official positioning is everyday development work, from fixing small bugs to building complete features. The same day, MiniMax reset Token Plan quotas for all users and issued an extra quota-reset card; from September 28 to October 7, check-in points in MiniMax Code are doubled for both new and existing users, and the points can be spent across any model.
Right now the model can only be reached through two doors: a Token Plan subscription, or MiniMax Code itself. A pay-as-you-go public API has not been opened yet.
What the docs confirm
MiniMax's open platform integration docs already list this model, and a few specs can be confirmed:
- A 1 million token context window, aimed at long documents, entire codebases and multi-step agent sessions;
- Claimed output speed of more than 100 tokens per second;
- Input support for text, images and video;
- An "effort" parameter that adjusts reasoning depth across five levels — low, medium, high, xhigh and max — defaulting to max if left unset.
One limitation is stated clearly: this model's reasoning cannot be turned off. If a call tries to disable it, the API returns a flat 400 error, and the docs suggest dialing effort down instead if you want lower latency. On integration, MiniMax offers both an Anthropic-style and an OpenAI-style endpoint, with the former marked as recommended.
What hasn't been announced
This launch was notably low-key. Tech media noted that MiniMax issued no formal announcement, no model card, no evaluation report and no pricing. The third-party leaderboard BenchLM had zero benchmark results logged for it as of September 27.
The same outlet also spotted a restricted repository on Hugging Face named MiniMax-M3.1-preview-private, roughly 250GB in size, which outsiders cannot download. Whether these are the Flash preview's weights, and whether it will eventually be open-sourced the way M3 was, MiniMax has not said. The docs also give no numbers on how much compute each of the five effort levels consumes or how much their latency differs.
A different release rhythm than M3
M3, released on June 1, played a different game: open weights, a 1 million token context and native multimodality all landed together, alongside MiniMax's self-developed sparse-attention architecture, MSA, and a full set of coding benchmarks — 59.0% on SWE-Bench Pro and 66.0% on Terminal-Bench 2.1. Back then, MiniMax made "affordable million-token context" its headline pitch, backed with plenty of numbers.
M3.1-Flash-Preview instead ships product first and documentation later: paying users get to use it inside MiniMax Code right away, while a quota reset and a check-in campaign pull people back in. The "Flash" and "Preview" in the name also signal where it sits — one word for a lightweight, speed-oriented tier, the other for something not yet finalized.
This pattern has become common among Chinese coding models lately. Xiaomi's MiMo-V2.6 splits into Pro and Flash tiers, with Flash taking the efficiency route; DeepSeek v4.1 Flash has spread through free quota handed out by the overseas coding tool OpenCode; Zhipu AI also has its GLM-5.3-Flash. Coding-tool users tend to care most about price and response speed, so shipping a lightweight tier first and filling in benchmarks later gets real-world feedback faster.
For developers in China wanting to try the model now, a Token Plan subscription or direct use of MiniMax Code is required. Effort defaults to the highest level, so for something as small as a bug fix, dialing it down a notch or two should mean a faster response and lower quota usage. How much faster it is than M3, and whether code quality suffers, will have to wait for MiniMax's own evaluations or third-party benchmarks.
Sources: MiniMax's official social media accounts, MiniMax open platform integration docs, AlphaSignal, CocoLoop, BenchLM; context length, effort tiers and integration limits are verified against MiniMax's documentation, M3 benchmark figures are self-reported by the vendor.