Ling 3.1 Flash is available through hosted APIs, with 560 billion total parameters and about 25 billion activated per token. Vercel’s September 30 release note offers a free promotion through October 13, 2026, and a 262K-token context window on its AI Gateway. That is immediate access to the model, not a release of its weights for local deployment.

The distinction matters because the word “Flash” and the smaller active-parameter count can make a very large mixture-of-experts model sound like a workstation model. Active parameters describe the subset used for each token’s computation. A server still needs access to the much larger collection of expert weights.

Hosted access has a dated free route

Vercel lists two identifiers: inclusionai/ling-3.1-flash and inclusionai/ling-3.1-flash-free. Its release note says the standard route is free during the promotion and begins billing afterward; the -free route stops serving rather than silently becoming a paid route. The free-route listing gives October 13 as the promotion end. A developer building a recurring workflow should choose the identifier intentionally and set a spend limit or a fallback before that date.

Novita’s own model page also offers serverless API access. It describes a one-million-token context in its summary, but the same page lists a 256K serverless context limit and a 32K maximum output in its supported-functionality section. Vercel states 262K for its route. The larger number should therefore not be treated as the context guaranteed by either public API route; providers can impose lower effective limits than a model-level claim.

Twenty-five billion active is not a 25B download

The 560B/25B description identifies a mixture-of-experts architecture. A rough lower-bound storage calculation for all 560 billion weights is 1.12 TB at two bytes per parameter, 560 GB at one byte per parameter, or 280 GB at four bits per parameter, before quantization metadata, embeddings, runtime buffers and key-value cache. These are arithmetic illustrations, not measured Ling 3.1 Flash memory requirements. The exact architecture, precision and deployment layout would be needed for a real hardware plan.

As of October 2, a search of InclusionAI’s Hugging Face model catalog does not show an official Ling 3.1 Flash weight repository. That does not prove weights will never be released or rule out private access. It does mean the currently documented way for an ordinary developer to try this version is through a provider, not by downloading the new checkpoint from the publisher’s model catalog. Ling 3.0 Flash is a separate, earlier model with its own published weights; its download and license do not transfer to 3.1.

What a useful evaluation can establish

The provider pages position Ling 3.1 Flash for coding, multi-step analysis and tool-using agents. They document model access and advertised architecture, but they do not establish task success for a specific workload. A fair trial should record the route and date, prompt and output lengths, context truncation, response latency, tool-call validity and retries. It should also compare the same tasks against the prior Ling 3.0 Flash and any other model under the same harness, with pricing checked again after the promotion.

For now, Ling 3.1 Flash is a sizable hosted reasoning model with a useful temporary trial path. Its practical limit is the provider route a developer can actually call, and its local-deployment story remains open until official 3.1 weights, license and deployment guidance appear.

Last Update: October 2, 2026