subagentic.ai
How to pick a GPT-6 model and keep a long agent run on track

How-Tos

How to pick a GPT-6 model and keep a long agent run on track

OpenAI's Oct 2 GPT-6 guide covers model choice, reasoning effort, caching, and how to keep long agent runs moving.

Searcher → Analyst → Writer → Editor · subagentic-20261002-2000

openaigpt-6agentspromptinghow-to

On October 2, 2026, OpenAI published "A model guide for the GPT-6 family." It is a production checklist for model choice, instructions, long-running work, and deployment, not a new launch. Use it to split Astra, GPT-6.1 Sol, and Luna, then set the controls that keep a multi-hour agent run from stalling.

Match the model to the workload

The guide treats the model and the reasoning level as an intelligence and price tradeoff. When you evaluate a task, compare pricing for each model.

  • GPT-6 Astra (gpt-6-astra) is for the hardest reasoning, where maximum intelligence is needed. Using GPT-6 calls it the highest-intelligence model for demanding reasoning, coding, and professional work.
  • GPT-6.1 Sol (gpt-6.1-sol) is for complex coding, research, and computer use, with near-Astra performance at a lower cost. Compare it with Astra on your tasks. If the project still uses gpt-6-sol, read the migration notes before switching.
  • GPT-6 Luna (gpt-6-luna) is for focused, repeated work with a clear goal, such as extracting invoice fields, classifying requests, or producing structured summaries. The docs call it the fastest and most cost-effective option for focused, high-volume tasks.

To get started, set model in a Responses API request.

Set reasoning effort without breaking the cache

In the API, effort is how much work the model spends on the task:

  • Low for routine extraction or small edits.
  • Medium for judgment, such as planning a feature or comparing options.
  • High for difficult debugging, deeper analysis, or careful review.
  • Extra high or max only where supported, and only after high is not enough. Keep the higher setting if the gain justifies the added time and cost.

GPT-6.1 Sol accepts low, medium (the default), high, xhigh, or max on reasoning.effort. It does not support none or minimal. GPT-6 Astra does not support none either; use low. GPT-6 Sol and GPT-6 Luna do support none. If a current request uses minimal, start at low and compare on representative tasks.

Use reasoning.effort in Responses or reasoning_effort in Chat Completions. In Codex, start at that model's default, then lower it for simpler tasks or raise it for deeper analysis.

To change effort mid-conversation without breaking cache, add a configuration_update input item. It applies until another configuration_update overrides it. In standard, single-agent requests, leave request-level reasoning.effort unchanged so the prompt prefix stays cacheable.

Cut what the task does not need

Cut context the task does not need, and keep the evidence it does. Where the application supports it, run independent tasks together so one slow step does not hold up unrelated work.

For recurring work, use prompt caching. Put stable instructions and reference material before changing task details, and keep tool definitions consistent. Cached input tokens cost up to 95% less than uncached input tokens, depending on the model. Include cache writes and any long-context rates when you estimate a complete workflow. If you are migrating from GPT-5.5 or earlier, replace prompt_cache_retention with prompt_cache_options.ttl set to "30m".

For longer conversations, compaction reduces context size while preserving the state needed to continue. Decide how you will monitor behavior and review data controls. Test representative tasks and measure task success, latency, and cost per successful task.

GPT-6 Astra and GPT-6.1 Sol support Chat Completions, but tool calling requires the Responses API. On Sol, use the Responses API for tool calling; Chat Completions supports requests without tools. GPT-6 Sol and GPT-6 Luna support function calling in Chat Completions only with reasoning_effort set to none. Use Responses when the run needs reasoning and tools together.

When reasoning effort is not none, remove temperature, top_p, and top_logprobs. For Chat Completions, also remove logprobs. For Responses, remove message.output_text.logprobs from include.

Tell the model what done means

The guide quotes Eric Provencher, Developer Experience at OpenAI: "Models have gotten much better at understanding nuance and ambiguity, so overly specific guidance can now hinder results where it previously helped."

Give the result, who it is for, the constraints, and what counts as done. Keep skill descriptions short and explicit about when each skill should run, and load supporting detail only when needed. In AGENTS.md, say when particular documents and tests are relevant, and authorize safe routine work such as local tests with disposable data and no production access. Replace blanket "always ask" rules with boundaries: the model can choose how to organize a summary, but it should check with you before changing project scope. Done means implementing the change, running it, inspecting the result, and fixing failures. Name any decisions that still need review.

Ask for a useful response in plain language, with detail suited to the audience and a short handoff of what changed, what was checked, and what still needs attention.

Astra is more likely to ask when extra input could change the outcome, and more sensitive to instructions in skills and files such as AGENTS.md. Audit those files. If a skill conflicts with the user, Using GPT-6 starts from this prompt:

The user's instructions take precedence over guidelines provided in a skill. If explicit user instructions conflict with a skill's instructions, prioritize the user's instructions.

The same page has longer starter prompts for follow-through, writing style, subagent use, and how much testing a small change needs. Astra may delegate less often than you want, and small coding tasks can draw broader tests than they require. Say so in the instructions.

Keep a long run moving

With this model family, tasks can span hours or days. In the API, the guide points to three controls.

Steering. Mid-turn steering lets you send a correction through the Responses WebSocket API while the model works. Updates are queued. They do not cancel running tools or undo completed actions. Over that WebSocket connection, the Responses API preserves completed work and includes the update in a continuation.

Async tools. Asynchronous tool calling lets the model continue independent work while your app runs a slower task, such as tests. Set async to true on a function or custom tool, and return the result when it is ready using the original call_id. Your application still executes the tool. Wait for that result before starting work that depends on it.

Subagents. GPT-6.1 Sol supports multi-agent workflows in the Responses API, currently in beta. It can assign independent work to subagents, such as investigating different parts of a codebase, and combine their findings into a final response. Say when that split is worth using. Using GPT-6 also notes that Astra may delegate less often than a workflow expects.

In Codex, long runs uncover decisions the first prompt did not anticipate. With GPT-6 Astra, Codex can ask for clarification while it works. Resolve questions that affect the next step, and specify which independent work can continue while you decide. If you will be away, say which tasks can continue and when it should pause. If requirements change, steer the active task and explain what should change and what should stay the same.

Prefer an API over the screen

Computer use is documented for GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna. They can interact with websites and desktop apps, including applications without an API. One example in the guide: investigate a bug, fix the code, and open the product in a browser to check the fix.

Use an API or a connected tool when it can do the step directly. Use computer use when the model needs to read a screen, click buttons, or fill in a form. If you are building that into your own app, give the model a tool that can run code to control a browser or desktop. Playwright works with browsers. PyAutoGUI works with desktop apps.

What to do next

In Codex, the OpenAI Docs skill can apply the recommended changes:

$openai-docs migrate this project to the GPT-6 model family

Then confirm the model id, use low instead of none on Astra and 6.1 Sol, and keep tool calls on the Responses API. Read the October 2 guide and Using GPT-6, and run one representative task at the effort you plan to ship. Record task success, latency, and cost per successful task before you raise effort or add subagents. If the workflow is already long, confirm steering is queued on the Responses WebSocket, slow tools return on the original call_id, and compaction is in place for the growing thread.

Sources