Skip to main content
Agentic retrieval uses an LLM to understand complex queries, decompose them into sub-queries, and intelligently aggregate results. While it has higher latency (2-5 seconds), it excels at nuanced questions that simple keyword or vector search can’t handle well.

How Agentic Retrieval Works

The agent:
  1. Analyzes the query to understand intent
  2. Decomposes into multiple sub-queries
  3. Executes each sub-query using hybrid search
  4. Aggregates and re-ranks results by relevance

When to Use Agentic Retrieval

Good Candidates for Agentic

Better Suited for Hybrid

Basic Usage

Set method to "agentic" on /api/v2/memory/search, and give the request a longer timeout than you would a hybrid search.
A search response carries episodes, profiles, agent_cases, agent_skills and unprocessed_messages. Which of these are populated depends on what the owner has and on include_profile.

Complex Query Examples

Example 1: Meeting Preparation

Example 2: Decision History

Example 3: Relationship Analysis

Cost and Latency Considerations

Agentic retrieval has higher resource usage:

Optimizing Agentic Queries

Fallback Strategy

Implement a tiered retrieval strategy:

Calling Agentic Retrieval from Async Code

The Python SDK is synchronous. For an async service, either run the client in a thread or call the endpoint directly with an async HTTP client.

Best Practices

Write detailed queries that explain what context you need:
Note that query cannot be empty. An empty string is rejected with a 422.
Set the timeout on the client, and always have a fallback:
The 1.x client does not retry on its own, so any retry or backoff is yours to add.
Cache results for repeated complex queries:
Reserve agentic for high-value queries where accuracy matters:

Next Steps

Concepts Guide

Compare all retrieval methods

Python Integration

Production patterns with timeout handling