Ant Group’s Ling-3.0-Flash Claims Trillion-Parameter Power From 5.1B Active Parameters
Content
Ant Group has released a model one-eighth the size of its flagship and told the market it barely matters. Ling-3.0-Flash, announced by the Hangzhou-based operator of Alipay and live on OpenRouter since July 23, carries 124 billion total parameters but activates just 5.1 billion per token — and, per the company’s launch thread, it “matches or beats our 1T flagship model on most benchmarks shown.” The Ant Group Ling-3.0-Flash release is not a capability announcement so much as a pricing signal: the company that spent late 2025 open-sourcing trillion-parameter models now argues the economics of AI agents will be decided by how little compute each action burns, not how much a model can theoretically do.
That argument carries extra weight coming from Ant. Three months before this release, the group’s blockchain arm shipped a platform for AI agents to transact in stablecoins at sub-cent scale. Cheap execution models are the other half of that equation — and Ant is now building both.
A Fraction Of The Flagship’s Compute, Most Of Its Score
The headline claim is aggressive. According to monitoring service figures circulating after the July 23 release, Ling-3.0-Flash beat the trillion-parameter Ling-2.6-1T on 11 of 12 internal benchmarks while activating roughly one-twelfth as many parameters per token. The press release frames the same result more conservatively, saying the model matches or surpasses industry leaders two to three times its parameter scale in foundational reasoning, instruction following, and long-context processing.
Both framings rest on Ant’s own testing. No technical report accompanied the launch, and no independent benchmark numbers existed at release — a departure from the previous generation, which shipped with a published Ling and Ring 2.6 technical report. Until third-party evaluations land, the parity claim is a well-constructed marketing position built on a real architecture, not a verified fact.
What is verifiable is the distribution strategy. The model launched simultaneously on OpenRouter and Vercel AI Gateway, plus several third-party inference platforms, with a free API window running through August 3, 2026. It supports both thinking and non-thinking modes, natively handles a 256K context window, and scales to one million tokens. After the free window closes, Ant says the weights will be open-sourced — more on that promise below.
How Ling-3.0-Flash Cuts The Inference Bill
The efficiency story rests on three architectural decisions, none of which involve simply making the model smaller.
The first is a native hybrid-linear attention design that stacks KDA and MLA layers at a 5:1 ratio. KDA — Kimi Delta Attention — is notable for its provenance: the mechanism originated in research published by Moonshot AI, the Beijing lab behind the Kimi models and one of Ant’s most direct rivals in the open-weight race. Ant’s version evolves its earlier Lightning Attention with fine-grained diagonal gating in the state-update rule, which the company says lets the model hold onto critical details across long documents and sprawling codebases. That a Chinese lab is shipping a competitor’s published attention mechanism at production scale says as much about the open-research flywheel among these labs as any benchmark chart.
The second is sparsity. Ling-3.0-Flash compresses its Mixture-of-Experts activation ratio from 1/32 in the previous generation to 1/64, halving the share of experts that fire per token. The third is serving infrastructure: a cluster-level hierarchical caching system that Ant says cuts time-to-first-token on long inputs by 60 to more than 80 percent by eliminating redundant computation across multi-turn conversations.
The company has been walking this road for a while. The Ling family debuted in October 2025 with the open-sourcing of Ling-1T and Ring-1T, the first open trillion-parameter reasoning model, then iterated through the Ling-2.5 generation in February and the Ling-2.6 series in April, each release trading raw scale for token efficiency. Ling-3.0-Flash is where the curve bends hardest: fewer active parameters than the April flash model, with a claim of flagship-tier output.
The Execution Node In A Planning-Execution Split
Ant is explicit that Ling-3.0-Flash is not trying to dethrone frontier reasoning models. The design target is what the company calls planning-execution separation: a large model plans, decomposes, and supervises, while a cheap, fast model carries out the high-frequency work — tool calls, code edits, searches, document processing. Ling-3.0-Flash is built to be that second model, a role Ant describes as a cost-controllable execution node.
The training regime reflects it. Ant says the model was refined across more than 10,000 interactive environments, with reinforced self-correction and long-horizon planning to keep multi-step tasks from drifting off course — the failure mode that quietly kills most production agent deployments. A companion multi-agent architecture lets separate agents divide labor and cross-check one another’s outputs, trading a little latency for fewer unilateral mistakes in high-frequency services.
The economics are the point. An agent that executes hundreds of tool calls per task lives or dies on per-token cost, and pricing history suggests where this lands: after its own free window closed, the predecessor Ling-2.6-flash settled at roughly $0.01 per million input tokens and $0.03 per million output tokens on OpenRouter — near the floor of the market. No platform has published post-window pricing for the 3.0 model yet, but Ant’s trajectory points the same direction.
Cheap Agents Are The Missing Half Of Ant’s Crypto Stack
For crypto markets, the interesting question is not how Ling-3.0-Flash scores on LiveCodeBench. It is why a payments conglomerate keeps driving the cost of machine cognition toward zero.
In April, Ant Digital Technologies — the group’s blockchain division — launched Anvita, a platform built for AI agents to hold assets, trade, and settle payments with minimal human involvement, completing sub-cent transactions instantly in USDC without invoices, subscriptions, or per-transaction human approval. Six months before that, the same division launched Jovay, an Ethereum layer-2 network designed for institutional tokenization of real-world assets, running a dual-proof architecture with built-in compliance checkpoints and, pointedly, no native token.
Those rails have a dependency problem: machine-to-machine payments measured in fractions of a cent only make commercial sense if the intelligence performing each action costs less than the action is worth. An agent paying $0.004 to query a data feed cannot be running on inference that costs ten times that per decision. A stable execution model priced in hundredths of a cent per million tokens is what makes the sub-cent economy arithmetic work.
Ant is not alone in seeing the connection — Visa, Coinbase, and Google are all building agentic payment infrastructure, and Coinbase shipped agent-native wallets earlier this year — but no Western competitor controls both a frontier-adjacent model lab and regulated settlement rails. Ant does, and its corporate boilerplate now lists AI and blockchain in the same breath. Whether Beijing’s continued hostility to retail crypto lets the group connect those pieces outside Hong Kong’s sandbox is the open strategic question; the technical pieces, at least, are converging fast.
Open Weights, On A Promise
One caveat deserves prominence. Ling-3.0-Flash is being marketed on the strength of Ant’s open-source track record, but as of this writing it is API-only: no weights, no model card, and no license file have appeared on the inclusionAI Hugging Face organization, which hosts roughly 150 of the lab’s earlier releases. The press release commits to open-sourcing the weights after the August 3 free window closes — a sequencing choice, plausibly, to concentrate a week of free usage on hosted channels before the download link goes live.
The lab’s history argues for taking the promise seriously; Ling-1T, Ling-2.6-flash, and Ling-2.6-1T all shipped under permissive licenses. But until the weights land, developers are testing a free trial, not adopting open infrastructure, and every benchmark figure remains internal. The gap between announcement and artifact is exactly the kind of detail that separates a genuine open-model release from a growth campaign wearing one’s clothes — and in ten days, Ant will show which this is.
The larger contest will take longer to settle. The frontier labs of 2024 competed on what a model could do; the labs shaping 2026 are competing on what an action costs. If agents genuinely become the dominant consumers of both inference and payments, the company that sets the floor price for a machine’s hour of work — and clears that machine’s invoices — collects a toll on everything above it. Ant has just cut its bid for the first half of that toll booth by an order of magnitude.