So Microsoft is routing its own products to cheaper in-house models whenever those models match frontier quality on a task.
Microsoft AI (MAI) models reportedly beat general-purpose frontier models on many tasks while using a fraction of the tokens.
The key is "Harness" again.
Substitutability, not capability, is what is actually being engineered here.
Harness, memory, context, and skills sit outside the weights by design, which turns the model into a replaceable component.
"The other key criteria to ensure that you are in control, is your evals should continue to hill climb even when any given model has been removed."
Your AI product should keep improving even if you replace or remove one underlying model.
Microsoft does not want Copilot's performance to depend entirely on OpenAI, Anthropic, or any single MAI model. The lasting advantage should come from the whole system around the model: product-specific evaluations, memory, tools, context, workflows, and user feedback.
Basically, looks l ike Microsoft is engineering a world where the model is the cheapest part to replace.