Snowflake Launches Dynamic Model Routing to Cut Enterprise AI cost
The company announced the capability within Cortex AI Gateway on August 18, alongside expanded access to open models.
Snowflake is introducing dynamic model routing to help enterprises reduce AI costs by automatically selecting the most suitable model for each task, as companies increasingly look beyond AI adoption metrics and focus on the economics of deploying AI at scale.
The company announced the capability within Cortex AI Gateway on August 18, alongside expanded access to open models. Snowflake said the new system is designed to improve what CEO Sridhar Ramaswamy calls “intelligence efficiency” — the ability to turn compute, models, data and context into measurable business impact.
"Customers can now define which models they approve and the tradeoffs they care about. From there, Cortex AI Gateway evaluates each task against those policies, along with real-world cost and performance data, and selects the best model for each use case," Ramaswamy said in a blog post.
Under the new routing system, enterprises can define which models they approve and the trade-offs they prioritise, including cost, performance and latency. Cortex AI Gateway then evaluates individual tasks against those policies and selects the model it determines is best suited for the job.
The approach reflects a growing challenge for enterprises deploying AI agents. While frontier models can deliver high performance, they can also be unnecessarily expensive for simpler workloads.
At the same time, the rapid improvement of open models is giving companies more options across different price and performance points. Ramaswamy said enterprises should therefore move away from the idea of standardising on one model.
“The question is moving from ‘Which model should we standardise on and where should we apply it?’ to ‘How do we continuously choose the best model for every task?’” he asked.
Snowflake said the routing system can also learn from the quality of model outputs. After one model completes a task, another model can evaluate the result, creating a feedback loop that can inform subsequent routing decisions.
The company said its initial benchmarks show that dynamic routing can provide better economics at a given quality level than relying on a single model. In one test involving agents building a dbt pipeline, Snowflake reported up to three times greater token efficiency while maintaining the same quality.
Another test showed a 25% improvement in token efficiency for engineering teams completing the same number of pull requests.
Snowflake is also expanding model choice on its platform. The company said it will add support for GLM-5.3 and DeepSeek-V4-Flash 0731, alongside models from providers including Anthropic, Google, Mistral AI, OpenAI and SpaceXAI.
The routing capabilities are being integrated into Snowflake CoCo and CoWork, allowing builders and business users to access different models, agents and workflows through a common interface.