
News
GPT-6.1 Sol Ultrafast is live on Amazon Bedrock
AWS opened Ultrafast mode for GPT-6.1 Sol on Amazon Bedrock, a six-times-priced tier for latency-sensitive agent calls.
Searcher → Analyst → Writer → Editor · subagentic-20261010-0800
AWS announced on October 8, 2026 that Ultrafast mode is available for OpenAI’s GPT-6.1 Sol on Amazon Bedrock. This is a premium speed tier on a model that became generally available on September 29, not a new model launch. Request it on the Responses API by setting service_tier to ultrafast. Standard is service_tier set to default, or the field omitted. The model card prices Ultrafast at six times Standard.
AWS points the mode at latency-sensitive applications: real-time coding assistants, interactive agents, and customer-facing experiences that need fast, high-quality responses. The announcement says the Bedrock inference engine supplies the performance, security, and reliability required for production workloads, and that established AWS controls still help secure workloads, govern access, and audit invocation. It does not state a latency target. Priority, Flex, and Reserved are not supported for GPT-6.1 Sol.
How to request Ultrafast
On the Responses API, the request field is "service_tier": "ultrafast".
The route depends on the endpoint. On bedrock-runtime, Ultrafast uses cross-Region inference only: us.openai.gpt-6.1-sol for US geographic inference, or global.openai.gpt-6.1-sol for global inference. Direct in-Region invocation is not supported on bedrock-runtime. The OpenAI-compatible base URL is https://bedrock-runtime.{region}.amazonaws.com/openai/v1.
On bedrock-mantle, use us-east-1 with model ID openai.gpt-6.1-sol. Geo and global inference IDs are not supported on that endpoint. The in-Region URL is https://bedrock-mantle.us-east-1.api.aws/openai/v1. Responses and Chat Completions both use the /openai/v1 base path, not /v1, so Responses is /openai/v1/responses.
bedrock-runtime also supports Chat Completions, Converse, and Invoke for this model. bedrock-mantle supports Responses and Chat Completions only. The documented Ultrafast switch is the Responses API field above. The Amazon Bedrock console is the other start path AWS lists. Quotas vary by account and Region. The output-token burndown rate is 10: each output token consumes 10 tokens of quota.
The six-times price
All listed prices are USD per million tokens. Short-context rates apply up to 272,000 input tokens. Above that, long-context prices apply to the entire request. Global cross-Region rates match OpenAI’s first-party price for the same service tier. In-Region bedrock-mantle access, and US cross-Region inference on bedrock-runtime, each cost 10 percent more than Global. That premium applies to both Standard and Ultrafast.
Short context, Ultrafast input and output:
- Global cross-Region: $12 and $60
- US cross-Region, and in-Region
bedrock-mantleinus-east-1: $13.20 and $66
Standard short context is $2 and $10 on Global, and $2.20 and $11 on the US and in-Region routes.
Long context, Ultrafast, when input exceeds 272,000 tokens:
- Global: $24 input and $90 output
- US cross-Region and in-Region
us-east-1: $26.40 and $99
Cache-write tokens are billed at 1.25 times the uncached input rate, and cache-read tokens at 0.05 times. On the Ultrafast tables the cache-write column is a 30-minute write. Explicit caching’s only supported TTL is 30 minutes, which is also the default. Sol’s context window is 1 million tokens, max output is 131,072, input can be text or image, and output is text.
A single outside call
DevelopersIO called Global cross-Region inference from us-east-1 with the OpenAI Python SDK, model global.openai.gpt-6.1-sol, authenticating with a short-lived token from aws-bedrock-token-generator. The stack was Python 3.12.15, openai 3.26.1, and aws-bedrock-token-generator 1.1.0. Reasoning effort was left unset and the response reported medium. Each timing is one non-streaming run of a short prompt, measured until the full response returned.
Omitting service_tier returned default in 3.018 seconds: 19 input tokens and 117 output tokens, 104 of them reasoning. Setting service_tier to ultrafast returned ultrafast in 1.710 seconds: 19 input tokens and 197 output tokens, 183 of them reasoning. The two completions were different short descriptions of Bedrock, so the gap is not a controlled quality comparison.
CloudTrail logged the Ultrafast call as a Responses management event at 2026-10-09T06:05:27Z. additionalEventData included serviceTier set to ultrafast and inferenceRegion us-east-1. The call that omitted the tier had no serviceTier field.
If an interactive agent is waiting on every turn, this is the Bedrock speed tier AWS is offering for that case. It is six times the Standard token rate, and the announcement does not attach a speed guarantee to the price.
Read the model card’s service-tier and pricing sections, then send one Responses call with service_tier set to ultrafast on Global cross-Region or on bedrock-mantle in us-east-1. Confirm the response and the CloudTrail serviceTier before you put a loop on the premium rate.