Article GPT-6 on Microsoft Foundry: A Guide to AI Agents

 

Article:

GPT-6 on Microsoft Foundry: Getting AI Agents into Production with Governance and ROI You Can Measure

 

 

By  Phil Hawkshaw / 9 Oct 2026  / Topics: Artificial Intelligence (AI)

Key Takeaways
  • Put Frontier AI into Production: GPT-6 Sol and Luna are now available on Microsoft Foundry, allowing businesses to build and deploy specialized, autonomous AI agents designed for complex tasks.
  • Control Your Costs and Governance: Move from speculative AI spending to clear financial discipline. Use Microsoft Foundry's built-in Model Router, token limits, and safety switches to prevent runaway costs from "agent sprawl."
  • Measure Real Business Value: Use the new ROI for Agents framework to connect agent costs directly to business results. This shifts the focus from "cost per token" to the more meaningful "cost per task."
  • Build on an Enterprise-Ready Foundation: Deploy these advanced AI models in secure EU Data Zones to meet GDPR and data residency requirements on a robust, expertly managed Azure foundation.

Earlier this year a CxO told me their AI spend had overtaken cloud as the fastest-growing line in their technology budget. Almost none of it had been planned. It was tokens: hundreds of millions of calls to frontier models, running in workflows nobody was really watching.

On 22 September, Microsoft made GPT-6 Sol and GPT-6 Luna generally available in Microsoft Foundry, joining GPT-6 Astra. It is a strong release, and it makes that conversation easier to have and easier to get wrong. Easier, because the price range across the three models is wide enough to change most bills. Easier to get wrong, because the price list is a small part of what decides whether an agent reaches production and stays there.

Insight Overview

What is Microsoft Foundry?

Foundry Agent Service is a managed platform for building, connecting, and scaling intelligent agents in an open, integrated, and enterprise-ready manner. It allows you to bring your preferred framework and model, deploying with a single command on a managed runtime with session isolation, native identity, and integrated observability.

Direct from Azure models, like GPT-6 Astra, are purchased and managed directly through Azure with a single license, consistent support, and no third-party dependencies. They include unified billing, governance, and PTU portability between models, all within Microsoft Foundry.

The Release: Three models, a hundredfold price gap

Microsoft positions Astra for demanding reasoning, software engineering and computer use, Sol for general production work, and Luna for high-volume extraction, summarisation and routing. All three are available as Standard deployments in the EU and US Data Zones as well as globally. These are the EU Data Zone prices:

Pricing Comparison

GPT-6 Astra vs. Sol vs. Luna

Metric GPT-6 Astra GPT-6 Sol GPT-6 Luna
Input, per million tokens $12.00 $2.40 $0.12
Output, per million tokens $60.00 $12.00 $0.60

Astra costs 100 times what Luna does, on input and on output. Sol sits at a fifth of Astra. As an illustration, take a step that reads 2,000 tokens and writes 300. At these prices it costs about $0.042 on Astra, $0.0084 on Sol and $0.0004 on Luna. Run it 100,000 times a month and the bill is roughly $4,200, $840 or $42.

I have seen the same principle work in our own engineering. We analysed roughly 90,000 UK Government contract documents on four on-premises GPUs. A frontier coding agent wrote the pipeline. Smaller models running on our own hardware did all of the reading. Data restrictions drove that design rather than cost, but it has the same shape: the expensive model does the thinking that needs it, and cheaper ones do the volume.

Cost per Task: The right measure, and an incomplete one

Microsoft's guidance is to look beyond price per token and understand cost per task. I agree with that, and I would go one step further, because the cost of a task has three parts: tokens, compute, and the human oversight around it. Tokens appear on the invoice. Oversight usually does not.

A cheaper model that sends more of its output to a person for correction can cost more per completed task than the dearer one, once that person's time is counted. Equally, a task done cheaply and correctly can still be a task nobody needed done. Cost per task tells you what a task costs. It does not tell you whether the client's outcome moved, and that is the question a CFO will eventually ask.

Answering it takes more than a good model router. It takes a structure around the agents.

The Structure: Four parts, and the model sits in one of them

At the centre of Insight's AI Centric Enterprise Architecture Framework is a simple picture, a stone gateway of the kind you see at Stonehenge: two upright pillars, a lintel laid across the top, and the ground all three stand on. Each part bears weight, and each has a different job.

Insight AI Centric Enterprise Architecture Framework

One pillar is technology: the estate agents run on, from infrastructure up to user experience. Foundry, Azure and Agent 365 sit here, and they are strong uprights. The other pillar is the business: the roles and the process, meaning who does the work, who is accountable for it, and how the work changes as AI takes on more of it. Governance is the lintel across both, deciding what gets built, who owns it, and what each capability is trusted to do. The foundation is knowledge: a formal description of what the organisation knows, which every agent reaches through to get to the data underneath.

In many of the conversations I have, the effort has gone into the technology pillar first. That is understandable, because it is where the demonstrations are. But most of what decides a production outcome sits in the other three. Four questions show where.

01

Who answers for each agent?

+

An agent is a role-holder, not a piece of technology. It has a job description, a position in the organisation and a set of responsibilities, whether it runs headless through a process or sits beside a person as a copilot. The question is whether the business declared those things or left whoever built the agent to infer them. A control plane can show you which agents exist and what they can reach. Naming the person who is accountable for each one is a business decision, and that accountability should stay human at every level of autonomy.

02

How much autonomy has each step earned?

+

It helps to place every AI-touching step on a five-point scale. At one end, no AI is involved. Next, the person does the work and uses AI as a tool. In the middle, AI performs some steps itself while a person still coordinates and resolves problems. Further along, AI coordinates a mix of AI and human workers. At the far end, people step in only when something breaks. The points differ in what the AI holds, not in who is accountable.

A step moves one point along only when business and technology readiness are evidenced together: demonstrated quality, an acceptable failure mode, an audit trail, tested escalation rules, and sign-off from the team that owns the outcome. Elapsed time is not a condition. A step can also move back, and that is the system working rather than failing. Foundry's evaluation and tracing supply much of the evidence. Deciding the evidence is enough remains a business call, and a good deal of pilot purgatory comes from nobody having defined what ready means.

03

What does the agent know?

+

Any competitor can buy GPT-6 on the same day, at the same price, as you. What they cannot buy is a formal description of what a customer means in your business, which rules govern a contract, and which source is authoritative for a price. An agent acts on whatever it retrieves. Where no source has been designated authoritative, it can retrieve a superseded definition and act on it with exactly the same confidence it would show if the definition were right. Models are bought. Understanding is built.

04

Where does the data sit, and who chose that?

+

Data residency is a good example of a decision that has a price. The EU Data Zone prices above run 20% higher than Global Standard for all three models. Where regulation such as GDPR, NIS2 or DORA requires processing in the EU, that is a cost worth carrying. Where the driver is a preference rather than a rule, it deserves a stated reason and a named owner. Treating every workload as mandated means paying a premium that nothing requires. Treating none as mandated means finding out about a residency requirement after the build. Our article on cloud sovereignty and compliance in Europe goes into the regulatory side.

What Insight Does: Starting from what is already running

The first conversation is rarely about a new model. It is about what is already running: agents a team built last quarter, AI features switched on inside software the organisation bought years ago, tools adopted without anyone being asked. We start by finding them, placing each step on that scale, and naming an owner.

From there, Insight supports the journey from experimentation to production in stages, working from our AI Centric Enterprise Architecture Framework and drawing on our Agentic Enterprise Accelerators. Business and technology readiness are checked together before anything advances:

  • Choosing models by evidence. We evaluate Astra, Sol and Luna against the organisation's own tasks in Foundry, and build a cost view for each use case that includes the human oversight.
  • Proving before scaling. New capability runs alongside the existing process first, so the comparison is made on the organisation's own cases before anything is handed over.
  • Agreeing what agents can rely on. Before agents are pointed at data, we work with the business to settle which sources are authoritative and how the organisation's own terms are defined. Our Agentic Enterprise Accelerators help build that shared description of the organisation, so agents are grounded in it.
  • Choosing what to build first. We help leadership rank use cases on value, deliverability and reuse, because the same capability is often requested by three functions in three different vocabularies and built three times. The accelerators support that prioritisation, and the governed delivery that follows it.
  • Governing the Microsoft estate. As a Microsoft Day 1 Launch Partner for Agent 365, Insight helps organisations set up and run the Azure, Foundry and Agent 365 foundation that all of this runs on, including region and residency choices.

None of it replaces what Microsoft has built. Foundry's Model Router, token limits and guardrails are the right controls for the technology pillar, and we configure them as such. The rest of the structure is what makes those controls mean something to the business. If you are still deciding where to begin, our guide on overcoming the AI implementation gap is a good place to start.

The Reframe: Same model, different outcome

GPT-6 will be available to your competitors on the same day, through the same platform, at the same price. The model will not be what separates one organisation's agents from another's. The owner, the evidence, the grounding and the choices about where data sits will.

If your agents are still running on the strength of the demonstration, that is a good place to start the conversation.

Headshot of Stream Author

Phil Hawkshaw

CTO EMEA, Insight