
Microsoft is planning to increase production of its forthcoming
semiconductor chip, "Maia 300," which is focused on inference.
"Maia 300" is Microsoft's internally designed, next-generation AI chip. Microsoft in January announced "Maia 200" as part of
the company's "heterogenous AI infrastructure" to serve multiple types of AI models, including the latest AI models from OpenAI.
Microsoft plans to release "Maia 300" in the fall,
reports The Information, citing an unnamed source.
Inference -- the process of drawing a logical conclusion, making a prediction, or uncovering hidden meaning based on existing evidence and
prior knowledge -- has been used for years in software to target ads and for other reasons.
Building inference into hardware such as chips makes it different and can increase
efficiencies.
In early 2026, an industry shift known as the "inference flip" occurred, including "economic and operational shifts where costs and computational power required to run AI
models have surpassed investments needed to train them."
advertisement
advertisement
This transition occurred as the industry moved from discovery and building larger models, to a utility phase focused on "thinking"
models.
In hardware, inference has become a computational process executed by special hardware such as Maia 300.
The company has touted Maia as part of a broader strategy
for continuous computational processing that supports each AI-driven function in Copilot and search.
The parts of Google's, Meta's, and Microsoft's businesses that are focused on hardware
consider this concept a "a major economic and technical turning point" for the tech industry where cost, energy, and hardware demand for running AI models overtake demand for building and training
them.
For now, inference with hardware is part of their solution to bring down costs.
This could make a substantial difference for advertisers. It could mean lower ad-serving costs and
lower costs to run hardware servers, accelerate real-time ad optimization, and power creative tools -- all because of inference capabilities in the chip, moving the industry away from expensive,
general-purpose hardware.
Microsoft is not the only company working on improving inference chips.
According to one report, Meta's new Iris chip is optimized to run AI tasks across Facebook, Instagram, and WhatsApp. It handles content ranking, ads, and
generative AI to free up third-party GPUs.
This is not Meta’s first move into custom silicon. The company has outlined its Meta Training and Inference Accelerator
(MTIA) roadmap, which includes several generations of AI chips for its own infrastructure.
Meta has not confirmed that Iris belongs to that family, but it follows the same strategy, according
to one report.
While
inference is a functional stage of modern AI, Meta CEO Mark Zuckerberg's vision of superintelligence represents a theoretical future that holds what he calls "superior creativity, strategy, and
problem-solving skills" that in his view should and will become available to all.
Zuckerberg says in his manifesto for AI superintelligence, published Monday, that AI will
better understand its user, including goals shared by the user, along with a list of many things that are of great importance to them.
According to his discussion, the agent will work on
the user's behalf to improve relationships, health, career, finances, home management, hobbies, and more. It will free up time for the things people enjoy, and help them accomplish more, but will also
offer privacy and security options, similar to how encryption works on WhatsApp.
Zuckerberg wants to make superintelligence affordable to all, and this is where the inference flip comes
in.
"For everyone to be part of [Zuckerberg's vision of the] future, everyone must have the ability to use superintelligence to improve their lives and shape the world," Zuckerberg
wrote. "We will offer free versions that will be accessible to billions of people."
As with search, revenue will be generated from those "who want to pay to use more compute [power], there
will be a dynamic auction mechanism that will guarantee that everyone gets the lowest price possible for the intelligence and compute they're using while also ensuring the capacity is used for
whatever people collectively find most valuable. This will ensure the benefits of superintelligence are distributed widely."
When he uses the phrase "dynamic auction mechanism," Zuckerberg
is trying to avoid the concentration of power.
In his stated strategy for monetizing Meta's future, "Personal Superintelligence" systems, he explained that while Meta will offer free basic AI
tools to billions of people, heavy compute users will be routed to this specialized market clearing model and auction.
The dynamic auction will work
like an automated and continuous real-time stock exchange or an ad auction, but for
computing power.
This is similar to algorithms that automatically match data centers trying to sell idle computing power with users who need to run heavy AI models.