Microsoft Develops Three In-House AI Models to Reduce Reliance on OpenAI

On April 2, Microsoft quietly launched three in-house AI models: MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2. There was no press conference or fanfare—they were simply made available on the Microsoft Foundry platform.

But the significance of this move goes far beyond the three models themselves.

What the Three Models Are

MAI-Transcribe-1: Speech-to-text. It supports the 25 most widely spoken languages globally and achieves a word error rate of 3.8% on the FLEURS benchmark, lower than OpenAI's Whisper and Google's Gemini. It is 2.5 times faster than Microsoft's existing Azure Speech service. Pricing is set at $0.36 per hour.

Notably, this model was built by a team of just 10 people.

MAI-Voice-1: Text-to-speech. It can generate 60 seconds of audio in under one second on a single GPU. It supports voice cloning from short audio samples. Pricing is $22 per million characters.

MAI-Image-2: Image generation. It has already climbed to third place on the Arena.ai text-to-image leaderboard and is at least twice as fast as the previous version. Pricing is $5 per million tokens (text input) and $33 per million tokens (image output). Early enterprise users include WPP, the world's largest advertising group.

Why This Matters

On the surface, Microsoft has released a few foundational models. In essence, it is building a safety net for itself.

Microsoft's relationship with OpenAI has not been a typical investor-investee dynamic since 2019. Microsoft injected $13 billion early on and secured exclusive rights to use OpenAI's technology to build products. Copilot and Azure AI are largely powered by OpenAI's core models.

But this heavy reliance is a double-edged sword. If OpenAI adjusts API pricing, changes terms, or is acquired or faces issues, Microsoft's AI product line would be vulnerable.

In September last year, Microsoft renegotiated its contract. While securing a $250 billion Azure cloud services commitment, it also obtained a key clause: allowing Microsoft to independently develop competitive AI models. This was previously prohibited.

The MAI team was established in November 2025, led by Microsoft AI CEO Mustafa Suleyman. Within five months of its formation, three models were launched.

The Real Moat Is Not Technology

Microsoft itself knows that these three models are not meant to technologically surpass OpenAI or Google.

The real trump card is distribution: The Microsoft Foundry platform hosts over 80,000 enterprise customers, covering 80% of the Fortune 500.

By directly offering "good enough, cheaper, and faster" built-in models to these 80,000 enterprises, Microsoft doesn't need to top every benchmark. As long as the models are adequate, stable, and integrate smoothly, enterprises will use them. This is Microsoft's natural advantage in the AI race—a distribution network that OpenAI and Anthropic lack.

Mustafa Suleyman's Next Move

Suleyman told the Financial Times that Microsoft's goal is to achieve "true self-sufficiency" in AI.

His timeline is 2027, when Microsoft plans to launch a frontier-level general-purpose LLM that will directly compete with OpenAI's flagship models.

From Microsoft's perspective, this is not about falling out with OpenAI but about building bargaining chips. After all, when you have your own models, the negotiation table looks very different come renewal time.

Sources: Microsoft takes on AI rivals with three new foundational models (TechCrunch); Microsoft launches three in-house AI models in direct challenge to OpenAI (The Next Web); Today we're announcing 3 new world class MAI models, CocoLoop, available in Foundry (Microsoft AI official blog)