Skip to content
Lumenta Digital
AI AutomationSeptember 15, 202610 min read
By Lumenta Digital

AI in 2026: What the New Models Actually Do for Your Business

Three flagship models landed in six weeks. Here is what changed, in plain English: which jobs these systems now finish on their own, and where the savings come from.

In the first week of September, Anthropic and OpenAI released new flagship models three days apart. Google had shipped one in August. If you run a business rather than a research lab, the honest reaction to that news is: so what?

It is a fair question, and most coverage answers it badly. Benchmark scores do not tell a business owner anything useful, and every release is described as a breakthrough, which means the word has stopped carrying information.

Something did change this year, though. Here is the plain version: what shifted, who released what, which jobs these systems now handle, and where the money comes from.

Three things changed, and only three

They can hold much more at once. The limit on how much a model can read in one go is called its context window. At Anthropic, OpenAI and Google it now sits at around a million tokens, or roughly 550,000 words. That is your employee handbook, three years of supplier contracts and a month of email in front of the system at the same time, with questions answered across the whole lot. Two years ago you fed these things a few pages at a time and stitched the answers together yourself.

They use software now, instead of writing about it. The phrase for this is computer use, and it is the year's biggest change. A model can open a system, find a record, fill a form, pull a figure out of a spreadsheet and move to the next step. Earlier versions could describe how to do that. These ones do it. That is the difference between an assistant who gives advice and one who does the filing.

The price per unit of work fell a long way. Vendors charge by the token, and on current models a token runs a little over half a word. Anthropic publishes $1 per million input tokens for its fastest model, so having a machine read around half a million words costs about a dollar. That figure matters more than any benchmark. It is what moves a task from "interesting demo" to "cheaper than doing it by hand."

Who released what

Anthropic: the Claude family

Anthropic now runs four models at once. That is where the cost control lives.

  • Claude Fable 5.1, released 1 September, is the heavyweight. It is built for long, multi-step work that has to hold together over hours, the kind of job where the system is still on task at step forty.
  • Claude Opus 5 handles complex business and coding work and is the sensible default for most demanding jobs.
  • Claude Sonnet 5 trades a little capability for speed and costs a fraction of the top tier. It is the one most production workflows actually run on.
  • Claude Haiku 4.5 is the fastest and cheapest, for high-volume, low-judgment work: sorting, tagging, extracting, routing.

You do not choose one of these and live with it. A properly built workflow sends each step to the cheapest model that can do it, and escalates only the steps that need the expensive one. Sorting a thousand emails does not need the heavyweight; drafting the reply to the one difficult email might. Businesses that skip this and route everything to the top model pay five to ten times more for the same result.

OpenAI: GPT-6 Astra

OpenAI released GPT-6 Astra on 3 September, describing it as its most capable model for business work, with stronger reasoning, better judgment in writing and design, and computer use at the centre of it. It has a context window above a million tokens.

OpenAI's own headline example is that Astra completed Financial Modeling World Cup challenges roughly four times faster than the human who won it. Treat vendor claims as vendor claims. The direction is clear enough: complicated spreadsheet work is now in range.

Astra is rolling out across ChatGPT Plus, Pro, Business and Enterprise, and through the API and the major cloud platforms. It follows a steady year of releases: GPT-5.4 in March, GPT-5.5 in April, and GPT-5.6 in three sizes in July.

Google: Gemini

Google made Gemini 3.7 Flash generally available in August, aimed at the workhorse tier: cheap and fast, built to run a great many small steps reliably rather than to win a reasoning contest. Its predecessor cut token consumption by up to 17%, which means the same job got cheaper.

The practical gap between these three vendors is narrower than the marketing implies. For almost every small and mid-size business, what you point the system at matters far more than which logo is on it.

What this makes possible that was not possible last year

The test that matters is which jobs in your building it can now finish without supervision.

  • Reading paperwork that is not tidy: phone photos taken at an angle, scans, handwritten notes, forms in layouts nobody standardised. This used to be the wall. It mostly is not.
  • Multi-step jobs across several systems: an enquiry arrives, a client record is created in the CRM, the intake email goes out, the call is booked, the file is opened. One instruction, five systems, no tabs.
  • Conversations, at any hour: a phone or web assistant that answers from your own approved material, books and reschedules, and passes the awkward ones to a person. The voices stopped sounding like robots about a year ago.
  • First drafts of judgment work: a quote assembled from a site-visit voice note and your price book. A proposal built from intake answers. Not finished work, and not meant to be. But the difference between a blank page and an eighty-percent draft is most of the afternoon.
  • Meetings that turn into actions: calls transcribed, summarised, and converted into tasks with owners, without anyone volunteering to take notes.

Where the money actually comes from

Four places. Only one is obvious.

The first is hours moved off the wrong people. Every business pays skilled staff to do unskilled work: renaming files, retyping invoices into a second system, chasing clients for documents. That hour is billed at the wrong rate or written off. Moving it is the largest and most reliable saving. It is also boring, which is why it gets overlooked.

The second is growth without proportional hiring, and it compounds. If volume rises 30% and headcount does not, the automation has paid for itself several times over and goes on doing so every year after that.

The third is speed converted into revenue. An enquiry answered at 9pm instead of 9am the next day is often the difference between a booked job and a competitor's booked job. It rarely appears in the business case and often turns out to be the biggest number in it.

The fourth is rework that never happens. Data entered once and passed between systems does not get transposed, duplicated, or filed against the wrong client. Finding and fixing those errors costs money, and almost nobody measures how much.

The model cost itself is usually the smallest line in the project at current published prices. The expensive parts are choosing the right process and wiring it into the systems you already run. Anyone selling you AI on the basis of model cost is discussing the wrong number.

To put figures against your own operation rather than someone else's, our savings calculator takes three inputs and estimates the hours and dollars in play.

The limits, stated plainly

These systems are confident when they are wrong. A model that has misread something does not sound any different from one that has read it correctly. Anything with real consequence, and anything a client will see, needs a person's approval. That is a design decision.

Automation inherits your process. If documents currently reach your business five different ways depending on which staff member the client knows, automating that produces five automated messes. Agreeing on one path in is part of the work, and it is the part that involves people rather than software.

A subscription is not a strategy. Buying seats and telling the team to use them produces a handful of enthusiasts and a majority who forget. The value sits in the workflows you build, not the licences you hold.

Some of it should wait. Not every capability demonstrated in a launch video is safe for a business your size to depend on. A good advisor will tell you which parts those are.

How Lumenta Digital puts this to work in your business

All of this only matters once it is pointed at a specific process in your business. That is the work we do.

We start by finding where your hours go. Our AI automation engagements open with a structured review of how your team works, to identify the one or two processes where automation pays off first. That is rarely the process people nominate. It is usually the repetitive, low-judgment work nobody thinks to mention, because it has always just been part of the day.

We prove one workflow before you commit to anything more. We build a single high-impact process on a defined scope, agree in advance what success looks like, and put it in front of your team to judge on results. You decide whether to extend it from what you can measure, not from what we promised. That is deliberate: it is where the risk comes out of the project. And if simpler automation would serve you better than anything sophisticated, we say so in the first call.

We connect the tools you already run: Outlook, QuickBooks, Xero, Excel, SharePoint, HubSpot, Pipedrive, Calendly, Teams. Your staff enter information once and it moves between your systems letter-perfect, instead of being retyped three times. We work inside your existing tools rather than asking you to migrate to ours.

Where the work needs judgment, we go a tier up. Advanced AI covers what rule-based automation cannot reach: agents that take a goal and work the steps with your own tools before handing you the decision; conversational and voice assistants that answer your phone and your website around the clock from knowledge you have approved; and vision and document AI that reads photos, scans and forms by their layout rather than just their text. Each one is proven on a single narrow capability first, measured against how you work today.

The guardrails are fixed. The AI proposes and a person approves anything that carries consequence, with you setting where that line sits. We build inside the accounts you control, agree in writing what data each workflow touches and where it is stored, and never feed your business or client data to public AI models for training. Our practices are built to respect PIPEDA and your own confidentiality obligations. If a capability stumbles, it falls back to your existing process rather than failing silently.

Some industries get a head start. We have built around the specific paperwork and pipelines of accounting firms and real estate, so those engagements begin with a good deal already understood. If you are in accounting, start with what AI automation really does in a firm rather than this article.

The first consultation is free, and it is a conversation about where your hours go before it is a conversation about tools. If the answer is that you do not need this yet, that is a perfectly good outcome of a phone call, and you will hear it from us.

Want to talk it through?

The first consultation is free, with no obligation.

Get a free quote