Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Note
Azure AI Search is available through the Azure portal, REST APIs, and Azure SDKs.
Important
Features, capabilities, or properties marked (preview) aren't covered by a service-level agreement, aren't recommended for production workloads, and might change or be constrained before they become generally available. The Azure AI Search preview terms apply to all preview functionality, whether it's standalone or part of a generally available feature.
In agentic retrieval, you can specify the level of large language model (LLM) processing for query planning and answer formulation. Use the retrieval reasoning effort (preview) to set LLM processing levels that affect costs and latency. Extra LLM processing improves relevance, but it also takes longer and uses billable LLM resources.
You can set this property in a knowledge base or a retrieve request. The knowledge base setting establishes the default for all queries, while the retrieve request setting overrides the default on a query-by-query basis.
Prerequisites
An existing knowledge base with at least one knowledge source and a model configuration.
Permissions to update knowledge bases. Configure keyless authentication with the Search Service Contributor role assigned to your user account (recommended) or use an API key.
If the knowledge base specifies an LLM, the search service must have a managed identity with Cognitive Services User permissions on the Azure AI services resource.
The 2026-08-01-preview REST API or an equivalent Azure SDK preview package: .NET | Java | JavaScript | Python
Choose a reasoning effort
Choose a reasoning effort based on the tradeoff you want between latency, cost, and retrieval depth.
Reasoning effort levels
| Level | Description | Recommendation | Limits |
|---|---|---|---|
minimal |
Disables LLM-based query planning to deliver the lowest cost and latency for agentic retrieval. It issues direct text and vector searches across the knowledge sources listed in the knowledge base, and returns the best-matching passages. Because all knowledge sources in the knowledge base are always searched and no query expansion is performed, behavior is predictable and easy to control. It also means the alwaysQueryKnowledgeSource property on a retrieve request is ignored. |
Use minimal for migrations from the Search API or when you want to manage query planning yourself. |
|
low |
The default mode of agentic retrieval, running a single pass of LLM-based query planning and knowledge source selection. The agentic retrieval engine generates subqueries and fans them out to the selected knowledge sources, then merges the results. You can enable answer synthesis to produce a grounded natural-language response with inline citations. | Use low when you want a balance between minimal latency and deeper processing. |
|
Set the reasoning effort in a knowledge base
This section demonstrates how to set the retrieval reasoning effort in an existing knowledge base. Although you can use this configuration for new knowledge bases, knowledge base creation is beyond the scope of this article.
To establish the default behavior, set retrievalReasoningEffort in the knowledge base definition.
### Set retrieval reasoning effort in a knowledge base
PUT {{search-url}}/knowledgebases/{{knowledge-base-name}}?api-version=2026-08-01-preview
Content-Type: application/json
api-key: {{api-key}}
{
"name": "{{knowledge-base-name}}",
"knowledgeSources": [ ... // OMITTED FOR BREVITY ],
"retrievalReasoningEffort": {
"kind": "low"
}
}
Reference: Knowledge Bases - Create or Update
Set the reasoning effort in a retrieve request
To override the default on a query-by-query basis, set retrievalReasoningEffort in the retrieve request body.
### Override retrieval reasoning effort in a retrieve request
POST {{search-url}}/knowledgebases/{{knowledge-base-name}}/retrieve?api-version=2026-08-01-preview
Content-Type: application/json
api-key: {{api-key}}
{
"messages": [ ... // OMITTED FOR BREVITY ],
"retrievalReasoningEffort": {
"kind": "low"
},
"outputMode": "answerSynthesis",
"maxRuntimeInSeconds": 30,
"maxOutputSize": 6000
}
Reference: Knowledge Retrieval - Retrieve