Models · 13 August 2026

On 13 August 2026, OpenAI and Cerebras announced GPT-5.6 Sol Ultrafast: up to 750 output tokens per second, which, on the figures put forward, is about 14 times the pace of the Standard offer (about 53 tokens per second). Access is in preview, by invitation. No public Ultrafast tariff and no general-availability date have been communicated. The press has cited, for Standard and Fast, $5 per million input tokens and $30 per million output. Cybernecs does not have access to this preview. What we can say to a VSE or SME is something else: speed changes how agents behave, and API cost is not an IT-department detail.

Why 750 tokens per second is not a demo gadget

A “slow” agent imposes a style: short replies, few round-trips, the user who drops off. A very fast agent allows loops: reread a ticket, query several tools, rephrase, check a document, come back to the human in a few seconds. The uses cited around the announcement (incident response, support, finance, e-commerce) are exactly those where latency kills the value. Support that answers after thirty seconds is not support. An incident copilot that arrives after the first human diagnosis arrives too late.

That speed has a structural price. At $30 per million output tokens on Standard/Fast, a chatty loop costs money. Ultrafast, if it stays in the same tariff ballpark (which nobody has published), will amplify token volume because teams will dare longer chains. The bill does not follow speed in a linear way: it follows the number of loops that speed makes acceptable.

What this means for a French director

First point: you probably do not have Ultrafast today, and you do not need it in order to decide. The 13 August announcement confirms a trend already visible: cloud providers are pushing latency down, on specialised infrastructure (here Cerebras). The centre of gravity remains American. The AI Act, in force on transparency since 2 August 2026, changes nothing there. Transparency is not sovereignty. Telling the user they are talking to AI does not say where the prompt goes.

Second point: a faster agent, wired to your emails, quotes or tickets, increases the surface of error per unit of time. A pricing hallucination, a document sent to the wrong recipient, a loop that exfiltrates too much context to the API. Agent cybersecurity is not a separate chapter. See our AI cybersecurity page.

Third point: field trades (site, workshop, reception) gain from short latency, provided the system holds up without depending on a cloud queue. That is the whole point of local execution for some flows, and of cloud for others. An AI agent for construction that must answer in a company car park, on an average connection, does not have the same specification as an internal HQ tool.

Five actions, without waiting for an OpenAI invitation

  • Measure the real cost of your current agents. Input/output tokens, tool calls, retries. A monthly table by use (support, quotes, HR). Without that counter, Ultrafast or not, you are flying blind.
  • Cap the loops. Max number of tool calls, max size of context sent, a ban on stacking endless rephrasings. Future speed will make these caps more necessary, not less.
  • Separate sensitive flows. Customer data, plans, payroll, security incident: ask the question of where it runs. An on-premise Synapse Box does not hit 750 tokens/s on a Cerebras cluster. It avoids sending the file into an API whose tariff and jurisdiction move without you.
  • Keep a human on committing acts. Price, contractual commitment, reply to a public incident: human validation, even if the draft arrives in a second. Article 50 already requires you to inform about the nature of the counterpart. Review, for its part, is your governance.
  • Do not tie your business to a preview. Invitation only, no public Ultrafast price, no GA date: this is not an architecture building block. It is a market signal.

Sovereignty: the alternative is not “as fast”, it is “controlled enough”

Comparing a French box to an invitation-only Ultrafast offer is a false debate if it is reduced to a throughput benchmark. The useful debate, for an SME, is: which flows need extreme latency, which flows need to stay inside your walls, which flows can bear an EU or US cloud with a clear contract. Cybernecs AI solutions for VSEs and SMEs are built on that separation, not on the race for tokens.

The AI Act does not ban GPT-5.6. It requires, since 2 August 2026, that the user know they are talking to AI, that synthetic content be marked (with a deadline of 2 December 2026 for machine-readable marking of systems already on the market), and that professional deepfakes be labelled. An ultra-fast agent that does not say who it is remains non-compliant, however impressive it is in a demo.

To frame an architecture choice

If you are hesitating between cloud API, France hosting and a local box, we set the comparison on your real flows (volume, data, hours, network), not on a keynote. Write to us via the contact page. We do not have Ultrafast. We have the habit of running agents that have to hold on Monday morning, not only on announcement day.

Sources: Cerebras, TechCrunch, OpenAI communications, Artificial Intelligence Regulation (AI Act).

Post a comment

Your email address will not be published.

Articles de la même catégorie