Skip to content

Ling-3.0-flash

Ling-3.0-flash is inclusionAI's instant (instruct) model, a 124B-parameter Mixture-of-Experts model with roughly 5.1B activated parameters per token. It is designed with token efficiency and production-scale agentic inference as key priorities, delivering fast responses and strong execution across coding, document processing, and lightweight agent workflows. Served free via Novita's own $0 pricing, with BYOK as a fallback.

Providers
Capabilities
AIHubMix
aihubmix-byok
$0
$0
—
Unavailable
OpenRouter
openrouter-byok
$0
$0
$0
Unavailable
Nous Research
nousresearch-byok
$0
$0
—
Unavailable
Usage analytics

Loading usage…

API & code
Uptime & Health
No uptime data yet

These providers haven't been health-probed for this model yet. The router still routes around upstreams that fail live requests — uptime fills in once probe history accrues.

Share cards
Ling-3.0-flash share card
Ling-3.0-flash
AIHubMix upstream share card
AIHubMix upstream
OpenRouter upstream share card
OpenRouter upstream
Nous Research upstream share card
Nous Research upstream
Nous Research upstream share card
Nous Research upstream
Credits
Use your own key

Run Ling-3.0-flash on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.

No BYOK keys configured for this model yet.

Share a key with the pool to earn credits for every request it serves, covering your plan cost.

Text generation
Context length262,144 tokens
Max output209,715 tokens
ArchitectureTransformer
Categorytext
ReleasedJul 20, 2026
Modalities
Capabilities