Beyond the prompt: AI models, reasoning, and speed
Understand model selection, reasoning effort, and speed in Claude and ChatGPT Work, and when to change settings to balance quality, time, and token use.
1. Introduction: Settings you may not have noticed
You open an AI chat, describe what you need, and wait for an answer. If the result is disappointing, rewriting the prompt is a natural next step. But the words you type are only part of the setup.
Claude and ChatGPT also offer controls that can change which AI model handles your request, how much analysis it performs, and how it carries out the work. These controls may sit beside the message box or inside a menu. You can use the product without ever opening them.
Understanding them matters because different tasks place different demands on AI. Turning an approved paragraph into a shorter email may benefit from quick exchanges. Checking whether a project plan is feasible requires examining dates, dependencies, and staffing together. The configuration suitable for one task may spend too many resources on the other, or provide too little analysis.
This article explains what those choices change and how to recognize when each becomes relevant. Examples focus on Claude’s chat application and ChatGPT Work. Availability varies by model, account, and interface; examples from coding tools are identified separately.
2. Model choice: The capabilities available for the task
The model is the trained AI system that interprets your request and generates a response. Changing models can change how reliably the system follows detailed instructions, handles ambiguity, relates information across documents, or completes a task involving several steps. Models also differ in response speed and resource requirements. These differences are the basis of the providers’ selection guidance. Anthropic’s model overview, OpenAI’s model guide.
Consider a project with two draft plans. Extracting explicitly stated deadlines into a table is relatively well defined: the model must locate each date and associate it with the correct task. A model designed for fast, repeatable work is a reasonable starting point.
Now ask whether the project can finish on time. Installation requires design approval, but approval is due on Wednesday and installation is scheduled for Tuesday. The installation team is also booked for testing on Thursday, and a supplier’s delivery is unconfirmed. The model must connect these facts and assess their consequences. That makes capability in planning and handling constraints more relevant than it was for extracting dates.
The distinction concerns what the model must do with the information. A short request can contain difficult trade-offs; a long document can contain straightforward, repetitive records. For model selection, examine the relationships, uncertainty, and decisions involved, alongside the volume of material.
How the current models fit into that distinction
The descriptions below summarize the providers’ positioning. The examples apply those descriptions to ordinary work; they are illustrative suggestions, not measured results or exclusive assignments. Several models may suit the same task.
| Model | Provider positioning | Illustrative use |
|---|---|---|
| Claude Haiku 4.5 | Emphasizes speed and low cost. | Sort incoming requests using categories you have already defined. |
| Claude Sonnet 5 | Balances capability and responsiveness across everyday work. | Turn meeting notes into an update that preserves decisions and unresolved questions. |
| Claude Opus 5 | Suits complex reasoning, planning, and technical work. | Assess a delivery plan against dependencies and limited staffing. |
| Claude Fable 5.1 | Targets especially demanding analysis and extended work. | Reconcile a policy with several amendments and document the remaining inconsistencies. |
| GPT-5.6 Luna | Suits clear, repeatable extraction and transformation. | Extract dates and owners from regularly formatted records. |
| GPT-5.6 Terra | Supports everyday reasoning and tool use. | Compare weekly reports and prepare a table explaining the changes. |
| GPT-5.6 Sol | Suits difficult, ambiguous assignments requiring judgment. | Compare proposals whose strengths conflict with different project priorities. |
| GPT-6 Astra | Targets complex work spanning several stages and tools. | Investigate supplier options and produce a supported comparison and spreadsheet. |
Sources: Anthropic’s model overview, Anthropic’s selection guidance, and OpenAI’s model guide. File creation and research also depend on the tools and source access available in the product.
When a task is clear and easy to check, starting with an efficient model can make sense. When it already involves difficult interpretation or consequential decisions, starting with a more capable model may be justified. Anthropic describes both approaches in its selection guidance.
Changing the model is one way to change the setup. Another is to keep that model and adjust the amount of analysis it applies.
3. Reasoning effort: Choosing how much analysis the task needs
After selecting a model, you may be able to adjust its reasoning effort. This setting influences how much work the model devotes to developing its response, including breaking a problem into parts, considering alternatives, and checking relationships between requirements. Increasing effort keeps the selected model in place.
The distinction becomes clearer in the project example. Extracting the installation date requires locating information already stated in the plan. Deciding whether installation can move to Thursday requires checking both the approval sequence and the team’s existing assignment. The second request involves consequences that need to be considered together.
Higher effort can support that additional analysis, but generally takes longer and uses more tokens—the units used to represent information processed or generated by the model. The benefit depends on whether the task needs the extra work. Section 6 explains token use in more detail.
Claude: Low, Medium, High, Extra high, and Max
In Claude, click the model name beside the send button to find the available effort settings. Anthropic groups Low and Medium as suitable for routine work, describes High as a general balance, and positions Extra high and Max for more demanding tasks.
The examples below translate that guidance into practical choices. They illustrate reasons to consider a level, rather than establish which level a task must use.
| Effort level | When to consider it | Example and selection rationale |
|---|---|---|
| Low | A narrowly defined task with explicit information and little uncertainty to resolve. | “Extract the task names, owners, and dates into a table.” The work primarily involves locating and organizing supplied facts. |
| Medium | Routine work that also requires selecting, relating, or organizing information. | “Turn these meeting notes into a project update covering decisions, actions, and unresolved questions.” The model must distinguish different kinds of information and assemble a coherent account. |
| High | Analysis involving dependencies, competing requirements, or conclusions that need justification. | “Check whether this delivery plan respects the approval sequence and staffing limits.” The model must evaluate the requirements together and explain any conflicts. |
| Extra high | Sustained technical work, particularly coding and tasks in which the AI carries out several steps using tools. | “Review this software migration plan across dependent services, identify failure paths, and develop a migration and recovery sequence.” The analysis must remain consistent across several connected technical decisions. |
| Max | Especially difficult reasoning where additional analysis is worth a longer wait and greater resource use. | “Evaluate whether the migration and recovery plan remains feasible when several services fail together.” The task requires examining interacting failure conditions and checking whether proposed responses introduce further conflicts. |
Low and Medium overlap in their intended use. The distinction in the examples is the amount of interpretation involved, not an official boundary between task categories. Similarly, choosing Max expresses a preference for greater effort; it does not establish that the answer will be correct.
Claude also exposes a Thinking setting on supported models. Thinking enables a process of working through the problem, while effort influences how much work Claude applies. In the Claude application, Thinking cannot be disabled for Fable 5.1 or Opus 5, although their effort remains adjustable.
ChatGPT Work: Light, Medium, High, Extra High, and Max
In ChatGPT Work, the control beneath the message box contains model and reasoning choices. Power presets combine settings, while Advanced exposes individual controls. The available options depend on the model and account.
OpenAI uses Light in ChatGPT Work for the setting called Low in the Codex command-line interface. Its guidance distinguishes narrowly scoped work, tasks requiring more planning, and difficult analysis involving multiple steps, sources, or trade-offs.
| Effort level | When to consider it | Example and selection rationale |
|---|---|---|
| Light | A focused request whose instructions and source information leave little to resolve. | “Convert this approved task list into a table without changing the content.” The request requires faithful transformation rather than a new planning decision. |
| Medium | Work that requires some planning and organization within clear constraints. | “Create a work sequence from these tasks, dependencies, and confirmed availability.” The model needs to connect the supplied information and organize a usable sequence. |
| High | A problem involving competing constraints or alternatives that need to be compared. | “Compare these two delivery plans against staffing limits and deadlines, and explain the trade-offs.” The model must assess how each option performs against several requirements. |
| Extra High | Complex work for which you are willing to allocate more analysis within the selected model. | “Extend the comparison to include a supplier delay and reduced staffing; identify which recommendations change and why.” The added scenarios create more relationships and consequences to examine. |
| Max | The hardest problems where depth takes priority over response time and usage. | “Develop and assess recovery options for several dependent projects facing simultaneous delays.” The task calls for sustained analysis of interacting constraints and possible responses. |
OpenAI groups High and Extra High within the same broad category of difficult work. The examples show increasing analytical demands, rather than a fixed cutoff between the settings. Max gives the selected model more time for the task. Ultra changes how work is distributed and is discussed in the next section.
Choosing a level from the work required
Start by identifying the operation you need: extract information, organize it, compare alternatives, or evaluate interacting consequences. A higher setting becomes more relevant as the request requires more planning, interpretation, and checking. Once the substantive analysis is approved, a narrow follow-up such as shortening the conclusion may justify returning to a lighter setting.
Effort levels are relative controls, not fixed allocations of seconds or tokens. Anthropic explicitly describes effort as a behavioral signal rather than a strict token budget. A difficult request can still receive analysis at a lower setting, and increasing effort does not guarantee a particular improvement.
Also distinguish the amount of analysis from the capabilities of the selected model. More effort is relevant when a suitable model needs to work through a more demanding problem. A change from straightforward extraction to complex technical judgment may also justify reconsidering the model itself. You do not need to exhaust every effort level before doing so.
The comparison method in Section 7 explains how to assess whether a change improves the result enough to justify its time and resource use.
4. Work modes and tools: What the system needs to do
Some requests require actions beyond analyzing the material already in the conversation. A tool gives the AI a way to perform an action, such as searching the web or working with a file. A product’s work mode organizes how it carries a task through those actions. The labels differ between products, so their purpose matters more than a shared word such as “research” or “advanced.”
Return to the unconfirmed supplier delivery in the project plan. Extra analysis can explain how that uncertainty affects the schedule. Establishing the delivery date requires evidence from the supplier, supplied by you or retrieved from an accessible source. That is a reason to obtain information, rather than simply increase effort.
Claude’s web search is useful for checking current facts. Research supports a broader investigation across multiple sources, including connected information where available. For example, checking one published product specification is narrower than investigating several suppliers and reconciling conflicting specifications. Anthropic’s guidance on search, thinking, and Research.
OpenAI distinguishes Chat, for questions and conversation, from Work, for carrying an assignment through to a reviewable result. Discussing the weaknesses in a project plan differs from asking for several files to be examined and a completed brief produced. The latter requires a workflow with access to the relevant materials and tools. OpenAI’s Chat and Work overview.
Some workflows can also distribute work. ChatGPT Work’s Ultra option assigns separate parts to additional AI agents working in parallel. OpenAI’s Max and Ultra descriptions. An agent is an AI worker assigned a subtask. Reviewing independent supplier files is an example of work that could be divided; the conclusions still need to be brought together.
These choices can complement one another: retrieve evidence, analyze it, and produce a document. Selecting a workflow should follow from those required actions. A task’s importance alone does not establish a need for research or parallel agents.
5. Speed: Understanding what makes a response faster
Speed becomes especially useful when waiting interrupts a series of short exchanges. During a live editing session, you may request a heading, shorten it, and adjust its tone. Once the factual content is settled, responsiveness can matter more than additional analysis.
There are different routes to a faster response. Selecting a model designed for responsiveness changes the capabilities available. Reducing effort asks the current model to spend less work on analysis. Some products also offer faster processing of a supported model at a higher price. These routes have different effects on quality and usage. Anthropic’s selection guidance, Claude Code’s comparison of speed and effort.
For the heading task, a responsive model or lighter effort is worth considering because the change is narrow and easy to inspect. For the project feasibility assessment, reducing analysis could remove work the task needs. If that assessment is urgent, a paid processing option may be relevant where supported; if it can run while you do other work, the shorter wait may have little value.
A separate example from coding tools: Claude Code offers Fast mode for supported Opus models. Anthropic describes it as retaining model capability with faster processing and higher per-token prices. For subscription users, it draws on usage credits, a paid balance separate from the included allowance. OpenAI also documents Fast mode for supported Codex models, with greater credit consumption. These examples explain why faster processing can cost more; they do not establish that the same controls appear in ordinary chat. Claude Code Fast mode, Codex speed options.
The practical choice depends on what waiting costs you and what the faster option changes. To assess the resource side, it helps to understand what the product counts.
6. Tokens and usage: What the answer consumes
A token is a unit of information processed by a model. In text, it may represent a word, part of a word, or another sequence of characters. There is no fixed one-word-to-one-token relationship. OpenAI’s explanation of tokens.
For a text task, usage can involve three components:
- Input: the instructions and material supplied to the model, potentially including conversation history, files, and results returned by tools.
- Visible output: the text produced for you, such as an answer, table, or report.
- Reasoning: additional tokens used for intermediate analysis on models that support it.
This is why answer length is an incomplete measure of usage. A one-paragraph recommendation based on several files can require substantial input and analysis. In OpenAI’s API—the interface software uses to call models—reasoning tokens are billed as output even though they are not displayed as the final answer. OpenAI’s reasoning-token accounting.
Effort also does not prescribe an identical token total for every request. Anthropic describes it as guidance for the model’s behavior. A difficult question can still require analysis at low effort. On Opus 5, lowering effort does not reliably shorten the visible response, so specify the length separately when you need, for example, a five-bullet summary. Claude’s effort documentation.
Capacity, allowances, and price measure different things
The context window is the model’s capacity for information within a request, with space also needed for generated content. A usage allowance limits activity over a period. Credits are a product’s accounting unit; the price or allowance charged for work depends on its rules. A request can fit within the context window while still using a substantial part of an allowance. Claude’s explanation of context and usage limits, OpenAI’s usage and pricing explanation.
Claude identifies model choice, conversation length, complexity, and features as factors affecting usage. ChatGPT Work and Codex use shared usage accounting; token-based credit rates vary by model and processing option. A token count therefore does not translate into the same number of credits or the same cost everywhere. Use the rules for the product and plan you are using. Claude usage limits, OpenAI pricing.
Resource management starts with a clear scope. For a review of delivery dependencies, identify the relevant plans and staffing constraints. Include information needed to interpret them, and specify the output: perhaps a table of conflicts with a short explanation. An unrestricted request to investigate every aspect of the project creates more work than that defined comparison.
For recurring tasks, consider the resources needed to reach an acceptable result, including corrections. A quick first response that repeatedly needs repair may provide less value than a slower response that meets the requirements. Which configuration achieves that must be established from actual use.
7. Applying the choices to one assignment
Imagine preparing the project brief from the two draft plans. The table brings the distinctions together around changes in that assignment. It identifies a reason to reconsider a setting; it does not prescribe a model switch for every stage.
| What the task now requires | Choice to reconsider | Why it matters |
|---|---|---|
| Extract dates and owners from explicit statements. | A responsive model with light analysis. | The work is primarily locating and organizing known facts. |
| Determine whether dependencies and staffing make the schedule feasible. | Model capability and reasoning effort. | The work now requires interpreting relationships and checking their combined consequences. |
| Resolve an unconfirmed delivery date. | Access to a relevant source or tool. | The missing element is evidence that cannot be established from the supplied plans. |
| Produce a complete brief with tables and source references. | A workflow that can handle the files and deliverable. | The assignment includes producing and checking an artifact. |
| Shorten an approved conclusion during a live review. | A faster configuration. | The substantive decision is settled, and the immediate task is a limited wording change. |
Before changing a setting, identify what the response needs to do better. A missing staffing limit calls for better input. An overlooked conflict, despite having the relevant evidence, gives you a reason to examine capability or effort. An answer that is correct but unnecessarily slow raises a different question about the value of its configuration.
To assess a change, keep the source material and requirements consistent and, where practical, adjust one setting at a time. For this project, check whether the result identifies overlapping commitments, respects dependencies, and flags unresolved evidence. Also record the wait, usage where visible, and corrections needed. For repeated work, include several representative cases rather than judging from one polished answer.
This turns selection into a decision about observable results. The settings provide ways to change the work; the requirements determine whether the change was useful.
8. Conclusion: Recognizing when the default needs reconsideration
You do not need to adjust the controls before every message. You need to recognize when the task changes: from extracting facts to interpreting constraints, from analyzing supplied information to gathering evidence, or from discussing a result to producing it.
Model choice, effort, tools, and processing speed address different parts of that work. Understanding those distinctions gives you a basis for choosing a configuration that meets the required standard with an acceptable wait and resource use.
Source note: Product-specific statements are based on the linked official OpenAI and Anthropic documentation, checked September 20, 2026. Availability, interface labels, and billing rules can change. The scenarios are illustrative, not performance-test results; selection advice is an editorial application of the documented distinctions.