Microsoft is quietly but decisively moving its image generation stack away from OpenAI models and toward its own MAI image family, starting with Bing Image Creator and now PowerPoint and Copilot experiences across Microsoft 365. This is not just a product refresh but a strategic pivot that reshapes how one of the largest software companies in the world plans to build and control its core AI infrastructure in the coming years. AI-driven cybersecurity strategies are increasingly important, highlighting the growing role of AI in various sectors.
Why this shift matters now
For the past several years, Microsoft’s generative imagery has been closely associated with DALL E, especially DALL E 3, through Bing Image Creator and Copilot features in Office. That integration symbolized a deep reliance on external frontier models at a time when most large platforms were still experimenting with how to bring generative media into mainstream products.
Today Microsoft is signaling that era is ending. MAI Image models are no longer experimental side options. They are being placed at the center of Bing, Copilot, and increasingly Office, making Microsoft’s own stack the default engine for many of the everyday image prompts sent by consumers and enterprises.
This matters because image generation is no longer a novelty embedded in one or two creative tools. It is becoming part of standard productivity workflows, branding pipelines, and presentation design. The decision about which model sits behind those experiences directly affects cost structures, compliance and data residency guarantees, creative quality, and how much leverage Microsoft has over partners that were once treated as indispensable.
Image generation is now core infrastructure, shaping costs, compliance, creative quality, and platform power
From DALL E to MAI: how we got here
When Bing Image Creator launched, it leaned heavily on OpenAI’s DALL E models as the primary text to image engine, reflecting Microsoft’s early strategy of hitching its generative roadmap to OpenAI’s rapid progress. At that stage Microsoft’s main differentiator was distribution rather than its own frontier models. The company focused on integrating OpenAI systems into Bing, Edge, and Microsoft 365 rather than building a full image stack internally.
The pivot began with MAI Image 1, introduced as Microsoft’s first fully internal text to image model and quietly rolled into Bing Image Creator and selected Copilot experiences as an option alongside DALL E 3 and GPT 4o. MAI Image 1 did not instantly replace DALL E, but it marked the moment when Microsoft started treating image generation as a first-party capability rather than something entirely outsourced.
The real inflection point came with MAI Image 2. Microsoft announced the second generation image model in March 2026, highlighting improved photorealism, more reliable text within images, and competitive performance that placed it in the top three on the Arena text to image leaderboard, just behind leading models from Google and OpenAI. At the same time, MAI Image 2 began rolling out across Copilot and Bing Image Creator and became available to enterprise customers via API on Microsoft Foundry, moving the model from lab curiosity to production workload engine.
What Microsoft has actually shipped
MAI Image 2 and MAI Image 2 Efficient
MAI Image 2 is a general-purpose text to image model that accepts natural language prompts and targets a wide range of styles, from photorealistic scenes to illustrations and commercial product imagery. It is accessible through the MAI Playground, Bing Image Creator, Copilot, and via API for organizations that need image generation at scale.
To address cost and latency, Microsoft introduced MAI Image 2 Efficient as a faster and more affordable variant designed for production scale use. In Azure AI Foundry, MAI Image 2 Efficient is priced at about 5 United States dollars per one million text input tokens and around 19.50 United States dollars per one million image output tokens, with optimizations that reduce output costs significantly while improving speed for high volume workloads. Microsoft and independent coverage describe this Efficient variant as a key part of the push to make MAI the default engine behind Copilot image generation, displacing DALL E in many experiences.
MAI Image 2.5 and MAI Image 2.5 Flash
Microsoft has already iterated beyond MAI Image 2. MAI Image 2.5 is an internally developed text to image and image editing model that accepts both prompts and reference images and debuted near the top of the Arena leaderboards for both generation and editing quality. It is positioned as a step change over MAI Image 2, with notable improvements in text rendering, stylized illustration, and commercial imagery that aim to match or exceed expectations set by earlier OpenAI systems.
MAI Image 2.5 is available to developers through Microsoft Foundry, with pricing around 5 United States dollars per one million text input tokens, 8 United States dollars per one million image input tokens, and 47 United States dollars per one million image output tokens, reflecting its more advanced capabilities and focus on premium quality. For organizations that need speed and lower cost at scale, Microsoft offers MAI Image 2.5 Flash, a variant tuned for fast generation and production workflows, with reduced pricing and latency compared to the base model.
Taken together, MAI Image 2, MAI Image 2 Efficient, MAI Image 2.5, and MAI Image 2.5 Flash form a tiered image stack that Microsoft can match to different product needs, from Bing and consumer creation to enterprise batch generation and high fidelity editing.
Bing Image Creator moves off OpenAI
Community reports and documentation around Bing Image Creator now describe a clear transition away from standalone DALL E 3 and toward Microsoft-controlled MAI and GPT-based image systems. In current interfaces, users see model menus that include MAI Image 2, GPT 4o, and in some regions DALL E 3, but the trajectory is to make MAI and native GPT image generation the main options while phasing out DALL E as OpenAI retires those API endpoints.
OpenAI scheduled the deprecation and removal of DALL E 3 from its public API for May 12, 2026, a decision that effectively ends support for the original DALL E 3 snapshots inside Microsoft tools that rely on those backends. This external timeline forces Microsoft either to adopt newer OpenAI image systems or to lean on its own MAI stack. The company’s choice has been consistent with its broader strategy: MAI models are becoming the default whenever Copilot is asked to generate an image, with DALL E now framed as a legacy option rather than a core engine.
This transition is not only technical. In many cases, unusual activity detection on Microsoft’s services can lead to browser challenges that ask users to confirm they are not robots and to review their security settings to address potential concerns. It changes how Microsoft can promise stability and compliance. With MAI models under its direct control, Microsoft can manage updates, training data policies, and geographical deployment in ways that are harder to guarantee when the core model is owned by a separate provider with its own roadmap and commercial constraints.
PowerPoint and Office join the shift
PowerPoint has been one of the most visible consumer-facing showcases for generative images inside Microsoft’s ecosystem, thanks to Designer and Copilot features that can generate slide imagery and layouts on demand. Historically, those features drew on OpenAI models such as DALL E 3, and some official Copilot documentation still references DALL E 3 as the image engine for certain experiences.
Recent reporting and product behavior indicate that this is now changing. From late July 2026 onward, multiple outlets describe PowerPoint Copilot features as being backed by MAI Image models rather than OpenAI systems, rerouting prompts toward MAI Image 2 and its variants for creative content across Office applications.
In practice, that means the same MAI stack that powers Bing Image Creator is increasingly responsible for presentation imagery, template suggestions, and visual assets inside Microsoft 365.
For users, the change may feel subtle at first. Slides still appear, and the Copilot interface looks familiar. Over time, however, the consistency of style, the handling of text in images, and the latency patterns will reflect the characteristics of MAI Image rather than DALL E. For organizations that demand predictable behavior across thousands of generated assets, those differences in speed, cost, and control become significant.
Strategic motives: cost, control, and competition
Microsoft’s move away from OpenAI and other external image models is best understood as part of a broader strategy to reduce dependence on external AI providers, manage costs, and regain end-to-end control over its AI stack. Public comments from Microsoft AI leadership have emphasized that the company pays substantial sums to partners such as Anthropic and OpenAI and that reducing and ultimately eliminating those recurring costs is a clear goal.
At the same time, Microsoft formally listed OpenAI as a competitor in its 2024 disclosures, underscoring a complex relationship in which OpenAI remains both a strategic partner and an entity that increasingly overlaps with Microsoft’s own product ambitions. By building and deploying MAI models at scale, Microsoft is signaling that it wants to compete directly on foundational models rather than simply resell partner capabilities.
MAI Image 2’s strong early performance on leaderboards and the rapid follow-up with MAI Image 2.5 give Microsoft credible technical footing to justify that shift, especially when combined with pricing structures that let the company tune cost and quality for different tiers of users. Once Microsoft can offer comparable or better quality with its own models, the economic case for continuing to depend heavily on external providers weakens.
Control is just as important as cost. Operating its own models allows Microsoft to align training, safety policies, and telemetry with its compliance and enterprise commitments, including region-specific data residency requirements that can be hard to guarantee when core model operations are managed by another company.
What this means for businesses and creators
For businesses that rely on Microsoft 365 and Azure, the shift to MAI Image models has several concrete implications.
First, the pricing and performance characteristics of MAI Image 2 Efficient and MAI Image 2.5 Flash provide clearer knobs to manage cost per asset for large-scale use. With transparent token-based pricing and options tuned for either fidelity or throughput, organizations can design workflows where high-priority assets use premium models while bulk generation runs on efficient variants.
Second, model consistency becomes more predictable. As Microsoft standardizes on MAI across Bing, Copilot, and Office, the same prompts are more likely to produce similar styles and behaviors across tools, simplifying brand control and reducing the need to retrain teams on slightly different model quirks in each product.
Third, vendor risk shifts. Because Microsoft will own both the platform and the model for much of its image generation stack, customers will increasingly depend on a single provider for both infrastructure and core AI capabilities. That can simplify procurement and compliance, but it also concentrates power and reduces the diversification that comes from using multiple unrelated model vendors.
For individual creators and knowledge workers, the change will mostly surface as improvements in speed, text fidelity, and editing capabilities. MAI Image 2.5’s integrated editing allows users to take an existing photo, specify adjustments, and receive modified versions that preserve identity and scene context, without needing separate tools or manual retouching. This supports a more iterative creative process within standard Office and Copilot flows.
Risks and open questions
Despite the clear strategic logic, there are real risks and uncertainties.
Quality parity with the very best OpenAI or Google image models is not guaranteed across all use cases. While MAI Image 2 and MAI Image 2.5 rank highly on public leaderboards, some niche domains and artistic styles may still favor specialized models, and it will take time for enterprises to validate MAI performance against their own datasets and visual standards.
There is also the question of concentration. As Microsoft reduces reliance on partners, it builds a more vertically integrated stack that can be highly efficient but may reduce competitive pressure in some segments. Customers that previously benefited from innovation across multiple vendors through Microsoft’s integrations might find themselves more limited to Microsoft’s own release cadence and priorities.
On the transparency front, not all details of the PowerPoint transition and internal deployment timelines are fully documented yet, and official Copilot materials still mention DALL E in some places. That gap between marketing and implementation can create confusion for customers trying to assess model provenance, especially in regulated industries where knowing exactly which model generated a given asset is important.
Finally, the broader ecosystem effects are worth watching. As Microsoft moves its own traffic off OpenAI image endpoints, the revenue mix and feedback flows for OpenAI’s image systems will shift, potentially impacting how quickly external models evolve in response to real-world usage. Similar dynamics apply to other partners such as Anthropic when their workloads are replaced by MAI equivalents.
Key takeaways and what to watch next
Microsoft is using MAI Image 2, its Efficient variant, and MAI Image 2.5 with the Flash option to replace OpenAI powered image generation in Bing Image Creator and PowerPoint, and to standardize Copilot image experiences around a first-party stack. The move is driven by a blend of cost control, technical confidence, and a desire to own foundational models rather than simply distribute them.
For organizations, the practical outcomes include more consistent behavior across Microsoft tools, clearer pricing structures for large-scale image generation, and a tighter coupling between the platform and the model. For the industry, the shift marks another step in the maturation of big platform providers from AI resellers to full-stack AI companies that build their own models and deploy them at scale.
Over the next year, several signals will be especially important.
- How quickly MAI Image fully displaces DALL E and other external models across all Copilot surfaces.
- Whether Microsoft can maintain or extend leaderboard performance as other providers iterate on their own image systems.
- How enterprises respond to the new pricing tiers and to the balance between MAI Image quality and cost.
- How clearly Microsoft documents model usage inside Office and Bing so that customers can trace provenance and comply with internal and external governance requirements.
The underlying trend is clear. Microsoft is shifting from distributing partner models to operating its own AI stack at scale, with MAI Image as a core pillar of that strategy. If the company can sustain quality and transparency while reducing dependence on external providers, this transition will mark a durable move toward self-reliance in its creative and productivity infrastructure.
Conclusion
Microsoft is quietly reshaping the foundation of its consumer AI experiences. Image generation in Bing and PowerPoint is no longer powered primarily by OpenAI models such as DALL E. Instead, Microsoft is routing those requests to its own MAI Image family, a set of in house models built and deployed by the Microsoft AI team. This shift matters because it changes who controls the cost, performance, and data pathways behind some of the most widely used AI features in productivity software today.
Why this move matters right now
For the past few years, Microsofts rapid AI push has leaned heavily on OpenAI for frontier capabilities, from the original Bing Chat launch to the earliest incarnations of Copilot. That partnership gave Microsoft a first mover advantage in generative AI, but it also created a deep operational dependency on another companys models and roadmap.
The decision to replace OpenAI image models with MAI Image inside Bing and PowerPoint signals that this dependency is starting to unwind in areas where Microsoft believes it can match or beat OpenAI on quality and cost. At a time when usage is exploding and GPU bills are rising fast, moving image workloads to internal models is not just a technical change. It is a business and strategic pivot that will shape how sustainable these AI features are for Microsoft and for customers who rely on them.
How we got here Microsoft and OpenAI in context
Microsofts relationship with OpenAI began as a cloud and investment partnership and evolved into an exclusive distribution arrangement for OpenAI models on Azure. Over several funding rounds, Microsoft committed more than ten billion dollars to OpenAI and positioned it as the frontier model partner for Copilot and Azure AI services.
That exclusivity started to loosen in late 2025. Microsoft and OpenAI renegotiated their agreement, ending Microsofts exclusive license to OpenAI models and allowing OpenAI to sell through other clouds such as Amazon Web Services. The revised arrangement kept OpenAI as the preferred frontier provider but gave Microsoft more freedom to build and deploy its own foundation models under the Microsoft AI brand.
Microsoft then built an internal team sometimes referred to as MAI Superintelligence, led by former DeepMind and Inflection AI co founder Mustafa Suleyman, with a mandate to create homegrown models across reasoning, code, speech, and vision. In early 2026 that team shipped its first three production models MAI Transcribe 1, MAI Voice 1, and MAI Image 2 all trained on Microsoft controlled data and compute rather than OpenAI infrastructure.
These initial launches were framed not as a break with OpenAI but as a move toward self reliance and a more balanced multi model strategy inside Copilot. OpenAI and Anthropic still handled the majority of Copilot requests, but MAI models began to appear in specific workflows where Microsoft could optimize for cost, latency, or data residency.
Inside the MAI Image family
MAI Image is Microsofts in house text to image line designed to compete directly with OpenAI and Google on coherence, instruction following, and text rendering fidelity. The core MAI Image 2 model entered public view in spring 2026 via Microsofts Foundry platform and the MAI Playground, targeting enterprise and developer use as well as integration into consumer products.
Shortly after the initial launch, Microsoft introduced MAI Image 2 Efficient, a more compact variant tuned for lower GPU usage and faster generation while preserving most of the visual quality of the flagship model. Microsoft has promoted MAI Image 2 Efficient as materially cheaper to run, citing pricing starting around five dollars per million input tokens and output token pricing that undercuts the flagship by roughly forty percent.
External benchmarks help explain why Microsoft is confident enough to replace OpenAI models in mainstream products. MAI Image 2 and its successor MAI Image 2 point five have appeared near the top of public leaderboards such as Arena, ranking close behind Googles Gemini image systems and OpenAI GPT Image variants on several text to image quality metrics. For Microsoft, this is evidence that its own image stack is no longer a research experiment but a production grade alternative that can stand next to the partners it previously depended on.
What is changing in Bing PowerPoint and Copilot
The most visible change is that MAI Image is now the default model when users ask Copilot or Bing to generate an image. Requests that previously flowed to OpenAI DALL E now route to MAI Image, which means Microsoft captures both the value and the infrastructure cost of serving those generations directly rather than paying licensing or usage fees to OpenAI.
Reporting indicates that MAI Image 2 is being rolled out across Bing Image Creator and PowerPoint Designer features, replacing OpenAI powered systems behind those interfaces. In parallel, Microsoft is swapping in other MAI models in Office contexts. Tens of thousands of Copilot prompts per week in Excel and Outlook are already being handled by MAI models rather than OpenAI or Anthropic equivalents.
This is part of a broader orchestration strategy. Microsoft is configuring Copilot to route different tasks to different models depending on scenario, cost, and compliance requirements. Internal MAI models now serve coding, transcription, voice, and image workloads, while OpenAI and Anthropic still run many of the complex reasoning and long form generation tasks. Rather than a single provider, Copilot is becoming a multi model fabric with Microsoft increasingly in control of the routing logic.
Strategic implications for Microsoft OpenAI and the wider ecosystem
From Microsofts perspective, the benefits of this shift fall into three buckets cost, control, and leverage.
1. Cost and efficiency
Image generation is computationally heavy, especially at consumer scale. By running MAI Image on its own infrastructure and tuning Efficient variants, Microsoft can reduce GPU consumption and cut inference costs per image, which is critical for keeping Copilot features economically viable as usage grows. Lower marginal costs also make it easier to offer generous free tiers or bundle AI features into Office and Windows without eroding margins.
2. Technical and data control
Owning the full image stack allows Microsoft to control when models are updated, how they are fine tuned for specific safety and brand guidelines, and where data is stored and processed. That matters for enterprise customers who care about data residency and regulatory compliance and for consumer facing services that must coordinate content filtering across regions and legal regimes. Internal models can also be more tightly integrated with Microsofts tooling, security, and telemetry pipelines.
3. Negotiating leverage with partners
Building competitive in house models gives Microsoft more bargaining power in future negotiations with OpenAI and other labs. If Microsoft can credibly say that its own models are good enough for many workloads, it can reserve OpenAI usage for truly frontier tasks and push for better commercial terms on those remaining dependencies.
For OpenAI, this transition is a reminder that even close strategic partners will eventually seek autonomy. Microsoft still describes OpenAI as its frontier model partner and continues to route many Copilot tasks to OpenAI systems, but the share of traffic handled by OpenAI is no longer guaranteed. Over time, if MAI reasoning models catch up to GPT level capabilities, more of that frontier traffic could migrate as well.
For the broader industry, Microsofts move validates an emerging pattern. Large platforms are increasingly pursuing multi provider strategies, combining investments in external frontier labs with internal model development to manage risk and cost. Microsofts combination of OpenAI, Anthropic, and MAI models inside Copilot mirrors similar hedging seen at other cloud and consumer companies.
What this means for businesses and everyday users
From a user experience perspective, most people will not notice that the backend model has changed unless they are watching technical release notes. The prompts and interfaces inside Bing and PowerPoint remain largely the same. The core questions are whether images arrive faster, whether they look better, and whether the system behaves more consistently across languages and edge cases.
Early data and leaderboard placements suggest that MAI Image is competitive with GPT Image and Gemini on visual quality, coherence, and text rendering, which are crucial for slide design and marketing collateral. If Microsofts cost claims hold up at production scale, enterprises could see improved performance without a price increase, or eventually more flexible pricing and usage caps.
For businesses that rely on Copilot and Office, the deeper significance is that Microsoft is taking direct responsibility for more of the AI stack. That can simplify procurement and compliance conversations, since customers can evaluate Microsofts own model documentation, training data policies, and safety frameworks rather than stitching together disclosures from multiple vendors. At the same time, companies that selected Microsoft partly because of its OpenAI partnership will need to track where and when OpenAI is still in the loop for critical workflows.
Risks unanswered questions and what to watch
There are real risks and uncertainties that accompany this pivot.
Quality drift and regression
When a model provider changes, even subtle differences in behavior can affect workflows that have been tuned around specific quirks. If MAI Image behaves differently than DALL E on certain niche prompts, design teams and automation pipelines might see unexpected outputs until they adjust. Microsoft will need robust monitoring and rollback strategies to catch regressions quickly.
Safety and content moderation
Microsoft now must enforce safety guardrails on its own models at scale, including for sensitive content and regional regulations. While that autonomy is a strategic advantage, it also increases the burden of maintaining consistent standards across rapid model iterations. Enterprises will want transparent documentation on how MAI Image is trained, filtered, and monitored.
Model sprawl and complexity
Copilot is evolving into a system that orchestrates multiple models across tasks and apps. That flexibility is powerful, but it can also create complexity and confusion for customers trying to understand which model is responsible for which behavior, and how to audit or constrain that behavior for governance purposes. Clear labeling, telemetry, and configuration options will be important to keep trust high.
Competitive pressure
Finally, Microsoft is stepping more directly into the arena as a model provider competing with OpenAI, Google, Anthropic, and rising startups. That competition can drive rapid improvement in MAI models, but it also risks spreading resources thin if Microsoft attempts to match rivals across every modality without clear prioritization.
The bottom line and what to look for next
The replacement of OpenAI image models with the MAI Image family in Bing and PowerPoint is a concrete milestone in Microsofts long term plan to own more of its AI stack while still partnering on frontier capabilities. It shows that internal models are no longer confined to niche or experimental features. They are becoming the default engines behind mainstream productivity experiences used by hundreds of millions of people.
Over the next year, several signals will reveal how successful this strategy really is. Watch for changes in Copilot pricing and quotas, which will reflect whether cost savings from MAI Image and other models are significant at scale. Track where Microsoft deploys new MAI versions inside Office, Windows, and Azure AI Foundry, and how often it highlights MAI rather than OpenAI in product marketing. Pay attention to feedback from designers, marketers, and everyday users on image quality and reliability, since their judgment will ultimately determine whether this internal pivot pays off.
If Microsoft can deliver equal or better quality at lower cost while preserving strong safety practices and transparent documentation, shifting image generation to MAI Image will look like a textbook example of a platform provider maturing its AI stack. If quality or trust wobbles, it will be a reminder that swapping out the engine beneath widely used tools is never just an internal technical change. It is a change that users will feel, and one that must be managed with care.








