> ## Documentation Index
> Fetch the complete documentation index at: https://docs.evermind.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Agentic Retrieval

> LLM-guided multi-round query expansion for complex questions

Agentic retrieval uses an LLM to understand complex queries, decompose them into sub-queries, and intelligently aggregate results. While it has higher latency (2-5 seconds), it excels at nuanced questions that simple keyword or vector search can't handle well.

## How Agentic Retrieval Works

```mermaid theme={null}
flowchart TD
    Q["User Query: 'What context would help me prepare for my meeting with Alice about the Q3 budget?'"]
    Q --> A["LLM Agent"]
    A --> S1["Sub-query 1:<br/>Alice profile preferences"]
    A --> S2["Sub-query 2:<br/>Q3 budget discussions"]
    A --> S3["Sub-query 3:<br/>Previous meetings"]
    S1 --> R1["Results 1"]
    S2 --> R2["Results 2"]
    S3 --> R3["Results 3"]
    R1 --> AGG["Aggregate & Rank"]
    R2 --> AGG
    R3 --> AGG
    AGG --> F["Final Results"]
```

The agent:

1. **Analyzes** the query to understand intent
2. **Decomposes** into multiple sub-queries
3. **Executes** each sub-query using hybrid search
4. **Aggregates** and re-ranks results by relevance

## When to Use Agentic Retrieval

| Use Agentic | Use Hybrid Instead |
| - | - |
| Complex, multi-part questions | Simple keyword lookups |
| "Prepare me context for..." | "Find messages about X" |
| Queries requiring reasoning | Direct topic searches |
| High-stakes retrieval | Real-time responses |
| Research and analysis | Chat applications |

### Good Candidates for Agentic

```python theme={null}
# Complex context gathering
"What context would help me understand the evolution of our pricing strategy?"

# Multi-faceted questions
"What are all the factors we discussed that affect the product launch timeline?"

# Relationship queries
"What connections exist between our customer feedback and the recent product changes?"

# Preparatory context
"Help me prepare for a conversation with the engineering team about technical debt"
```

### Better Suited for Hybrid

```python theme={null}
# Simple lookups
"What did Alice say about the budget?"

# Direct searches
"Find discussions about Kubernetes"

# Recent context
"What did we discuss in yesterday's meeting?"
```

## Basic Usage

Set `method` to `"agentic"` on `/api/v2/memory/search`, and give the request a longer timeout than you would a hybrid search.

<CodeGroup>
  ```python Python SDK theme={null}
  from everos_cloud import EverOS

  # Agentic retrieval takes seconds, so raise the client timeout
  client = EverOS(api_key=EVEROS_API_KEY, timeout=60.0)

  def agentic_search(user_id: str, query: str, top_k: int = 10):
      """Perform agentic retrieval."""
      return client.search(
          query,
          user_id=user_id,          # exactly one of user_id / agent_id is required
          method="agentic",
          top_k=top_k,
          include_profile=True,     # bring the user's profile into the result
      )

  result = agentic_search(
      user_id="user_alice",
      query="What context would help me prepare for discussing the product roadmap with stakeholders?",
  )

  print(f"Found {len(result.episodes)} relevant episodes")
  for ep in result.episodes:
      print(f"- [{ep.score:.3f}] {ep.summary[:100]}...")
  ```

  ```python HTTP theme={null}
  import requests

  BASE_URL = "https://api.evermind.ai"
  headers = {
      "Authorization": f"Bearer {EVEROS_API_KEY}",
      "Content-Type": "application/json",
  }

  def agentic_search(user_id: str, query: str, top_k: int = 10) -> dict:
      """Perform agentic retrieval."""
      payload = {
          "user_id": user_id,
          "query": query,
          "method": "agentic",
          "top_k": top_k,
          "include_profile": True,
      }

      response = requests.post(
          f"{BASE_URL}/api/v2/memory/search",
          json=payload,
          headers=headers,
          timeout=60,  # longer timeout for agentic
      )
      response.raise_for_status()
      return response.json()["data"]

  data = agentic_search(
      user_id="user_alice",
      query="What context would help me prepare for discussing the product roadmap with stakeholders?",
  )

  episodes = data["episodes"]
  print(f"Found {len(episodes)} relevant episodes")
  for ep in episodes:
      print(f"- [{ep['score']:.3f}] {ep['summary'][:100]}...")
  ```
</CodeGroup>

<Note>
  A search response carries `episodes`, `profiles`, `agent_cases`, `agent_skills` and `unprocessed_messages`. Which of these are populated depends on what the owner has and on `include_profile`.
</Note>

## Complex Query Examples

### Example 1: Meeting Preparation

```python theme={null}
def prepare_meeting_context(user_id: str, meeting_topic: str, attendees: list):
    """Gather comprehensive context for a meeting."""

    attendee_names = ", ".join(attendees)
    query = f"""What context would help me prepare for a meeting about {meeting_topic}?
    Attendees include: {attendee_names}.
    I need:
    - Previous discussions on this topic
    - Relevant decisions and outcomes
    - Any concerns or blockers mentioned
    - Attendee preferences and working styles"""

    return client.search(
        query,
        user_id=user_id,
        method="agentic",
        top_k=15,
        include_profile=True,
    )

context = prepare_meeting_context(
    user_id="user_alice",
    meeting_topic="Q3 product roadmap",
    attendees=["Bob (Engineering)", "Carol (Product)", "Dave (Sales)"],
)
```

### Example 2: Decision History

```python theme={null}
def trace_decision_history(user_id: str, decision_topic: str):
    """Trace the evolution of decisions on a topic."""

    query = f"""Trace the history of decisions and discussions about {decision_topic}.
    I want to understand:
    - What options were considered
    - What factors influenced the decisions
    - Who was involved in the discussions
    - What the outcomes were
    - Any changes or reversals over time"""

    return client.search(
        query,
        user_id=user_id,
        method="agentic",
        top_k=20,
    )

history = trace_decision_history(
    user_id="user_alice",
    decision_topic="choosing our cloud provider",
)
```

### Example 3: Relationship Analysis

```python theme={null}
def analyze_topic_relationships(user_id: str, topics: list):
    """Find connections between multiple topics."""

    topics_str = ", ".join(topics)
    query = f"""Find connections and relationships between these topics: {topics_str}.
    Look for:
    - How these topics have been discussed together
    - Dependencies or conflicts between them
    - People involved in multiple topics
    - Timeline overlaps"""

    return client.search(
        query,
        user_id=user_id,
        method="agentic",
        top_k=15,
    )

relationships = analyze_topic_relationships(
    user_id="user_alice",
    topics=["customer feedback", "product features", "technical debt"],
)
```

## Cost and Latency Considerations

Agentic retrieval has higher resource usage:

| Metric | Hybrid | Agentic |
| - | - | - |
| Latency | 200-600ms | 2-5s |
| LLM Calls | 0 | 1-3 |
| Search Operations | 1 | 3-5 |

### Optimizing Agentic Queries

```python theme={null}
# 1. Use appropriate top_k - don't over-request
#    -1 (the default) lets the engine decide; explicit values are 1 to 100
client.search(query, user_id=user_id, method="agentic", top_k=10)

# 2. Scope to one owner - exactly one of user_id / agent_id is required
client.search(query, user_id="user_alice", method="agentic")
client.search(query, agent_id="support-bot", method="agentic")

# 3. Narrow the search space with app_id / project_id
#    Reads must use the same pair the memories were written under
client.search(query, user_id=user_id, method="agentic",
              app_id="support-bot", project_id="eu-west")

# 4. min_score is honoured by the episode hybrid path only; agentic ignores it.
#    Filter the returned scores yourself, or run hybrid with a floor first.
hits = client.search(query, user_id=user_id, method="agentic", top_k=10)
episodes = [e for e in (hits.episodes or []) if e.score >= 0.3]
```

## Fallback Strategy

Implement a tiered retrieval strategy:

```python theme={null}
from everos_cloud import EverOS, EverOSAPIError

fast_client = EverOS(api_key=EVEROS_API_KEY, timeout=10.0)
slow_client = EverOS(api_key=EVEROS_API_KEY, timeout=60.0)


def tiered_retrieval(user_id: str, query: str, complexity: str = "auto"):
    """Use appropriate retrieval based on query complexity."""

    if complexity == "auto":
        complexity = estimate_complexity(query)

    if complexity == "simple":
        return fast_client.search(query, user_id=user_id, method="hybrid", top_k=10)

    try:
        return slow_client.search(query, user_id=user_id, method="agentic", top_k=10)
    except (EverOSAPIError, TimeoutError):
        # Fall back to hybrid if agentic is unavailable or too slow
        return fast_client.search(query, user_id=user_id, method="hybrid", top_k=10)


def estimate_complexity(query: str) -> str:
    """Estimate if a query needs agentic retrieval."""
    complex_indicators = [
        "prepare", "context", "help me understand", "trace",
        "relationship", "connection", "evolution", "history of",
        "factors", "all the", "comprehensive",
    ]

    query_lower = query.lower()

    if any(indicator in query_lower for indicator in complex_indicators):
        return "complex"

    if query.count("?") > 1 or " and " in query_lower:
        return "complex"

    return "simple"
```

## Calling Agentic Retrieval from Async Code

The Python SDK is synchronous. For an async service, either run the client in a thread or call the endpoint directly with an async HTTP client.

<CodeGroup>
  ```python Run the SDK in a thread theme={null}
  import asyncio
  from everos_cloud import EverOS, EverOSAPIError

  client = EverOS(api_key=EVEROS_API_KEY, timeout=60.0)


  async def agentic_search(user_id: str, query: str, top_k: int = 10):
      """Await the synchronous client without blocking the event loop."""
      return await asyncio.to_thread(
          client.search, query,
          user_id=user_id, method="agentic", top_k=top_k,
      )


  async def search_with_fallback(user_id: str, query: str, prefer_agentic: bool = True):
      if prefer_agentic:
          try:
              return await agentic_search(user_id, query)
          except (EverOSAPIError, TimeoutError):
              pass  # fall through to hybrid

      return await asyncio.to_thread(
          client.search, query,
          user_id=user_id, method="hybrid", top_k=10,
      )
  ```

  ```python Async HTTP client theme={null}
  import aiohttp

  BASE_URL = "https://api.evermind.ai"


  class AgenticEverOSClient:
      """Client optimized for agentic retrieval."""

      def __init__(self, api_key: str, base_url: str = BASE_URL):
          self.base_url = base_url
          self.headers = {
              "Authorization": f"Bearer {api_key}",
              "Content-Type": "application/json",
          }

      async def _search(self, payload: dict, timeout: int) -> dict:
          client_timeout = aiohttp.ClientTimeout(total=timeout)
          async with aiohttp.ClientSession(timeout=client_timeout) as session:
              async with session.post(
                  f"{self.base_url}/api/v2/memory/search",
                  json=payload,
                  headers=self.headers,
              ) as response:
                  response.raise_for_status()
                  body = await response.json()
                  return body["data"]

      async def agentic_search(self, user_id: str, query: str,
                               top_k: int = 10, timeout: int = 60) -> dict:
          return await self._search(
              {"user_id": user_id, "query": query, "method": "agentic",
               "top_k": top_k, "include_profile": True},
              timeout,
          )

      async def search_with_fallback(self, user_id: str, query: str,
                                     prefer_agentic: bool = True) -> dict:
          if prefer_agentic:
              try:
                  return await self.agentic_search(user_id, query, timeout=60)
              except (asyncio.TimeoutError, aiohttp.ClientError):
                  pass  # fall through to hybrid

          return await self._search(
              {"user_id": user_id, "query": query, "method": "hybrid",
               "top_k": 10, "include_profile": True},
              10,
          )


  async def main():
      client = AgenticEverOSClient(api_key=EVEROS_API_KEY)

      data = await client.search_with_fallback(
          user_id="user_alice",
          query="What context do I need to understand our pricing strategy evolution?",
          prefer_agentic=True,
      )

      print(f"Found {len(data['episodes'])} episodes")

  asyncio.run(main())
  ```
</CodeGroup>

## Best Practices

<AccordionGroup>
  <Accordion title="Query Formulation">
    Write detailed queries that explain what context you need:

    ```python theme={null}
    # Good: Detailed, explains intent
    query = """What context would help me prepare for discussing
    technical debt with the engineering team? I need to understand
    past discussions, proposed solutions, and any blockers mentioned."""

    # Bad: Too vague
    query = "technical debt"
    ```

    Note that `query` cannot be empty. An empty string is rejected with a 422.
  </Accordion>

  <Accordion title="Timeout Handling">
    Set the timeout on the client, and always have a fallback:

    ```python theme={null}
    # Longer timeout for agentic
    client = EverOS(api_key=EVEROS_API_KEY, timeout=60.0)

    try:
        result = client.search(query, user_id=user_id, method="agentic")
    except (EverOSAPIError, TimeoutError):
        result = client.search(query, user_id=user_id, method="hybrid")
    ```

    The 1.x client does not retry on its own, so any retry or backoff is yours to add.
  </Accordion>

  <Accordion title="Result Caching">
    Cache results for repeated complex queries:

    ```python theme={null}
    from functools import lru_cache
    import hashlib

    def cache_key(user_id: str, query: str) -> str:
        return hashlib.md5(f"{user_id}:{query}".encode()).hexdigest()

    # Cache expensive agentic results
    @lru_cache(maxsize=100)
    def cached_agentic_search(cache_key: str, user_id: str, query: str):
        return agentic_search(user_id, query)
    ```
  </Accordion>

  <Accordion title="Selective Use">
    Reserve agentic for high-value queries where accuracy matters:

    ```python theme={null}
    # Use agentic for:
    - User explicitly asks for comprehensive context
    - Preparing for important meetings/decisions
    - Research and analysis tasks

    # Use hybrid for:
    - Real-time chat responses
    - Simple lookups
    - Frequently repeated queries
    ```
  </Accordion>
</AccordionGroup>

## Next Steps

<CardGroup cols={2}>
  <Card title="Concepts Guide" icon="brain" href="/cloud/concepts/memory-types-retrieval">
    Compare all retrieval methods
  </Card>

  <Card title="Python Integration" icon="python" href="/cookbook/python-integration">
    Production patterns with timeout handling
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.