> ## Documentation Index
> Fetch the complete documentation index at: https://docs.evermind.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Personal AI Assistant

> Build a 1:1 AI assistant that remembers user preferences and context across sessions

Build an AI assistant that remembers user preferences, past conversations, and
context, and carries that memory across sessions. This guide integrates EverOS
into a personal assistant loop: recall relevant memory, generate a grounded
reply, then persist the turn.

## Architecture

The assistant runs the same loop every turn:

1. **Recall:** search the user's memory for context relevant to the message
2. **Generate:** call your LLM with that context
3. **Persist:** store the user and assistant messages back into EverOS

Extraction and consolidation happen in the background. You never manage them.

## Setup

```bash theme={null}
pip install everos-cloud
export EVEROS_API_KEY="your_api_key"
```

```python theme={null}
from everos_cloud import EverOS

client = EverOS()  # reads EVEROS_API_KEY
```

## Store conversation turns

Store each turn into a session. Each message's `sender_id` is who it's attributed
to: the user's id for their messages, `"assistant"` for replies. Writes are
asynchronous by default; extraction runs in the background.

```python theme={null}
def store_turn(session_id: str, user_id: str, user_msg: str, assistant_msg: str):
    client.add(
        session_id=session_id,
        messages=[
            {"sender_id": user_id, "role": "user", "content": user_msg},
            {"sender_id": "assistant", "role": "assistant", "content": assistant_msg},
        ],
    )
```

## Retrieve relevant context

Before generating, search the user's memory. Episodes are narrative summaries of
past conversations and are available within seconds of extraction. The
**profile** (consolidated preferences and traits) builds up over time in the
background; request it with `include_profile=True` and use it when present.

```python theme={null}
def get_context(user_id: str, session_id: str, user_message: str) -> str:
    results = client.search(
        user_message,
        user_id=user_id,
        method="hybrid",
        top_k=5,
        include_profile=True,
        filters={"session_id": session_id},   # also return this session's live tail
    )

    parts = []

    # Consolidated user profile (fills in over time via background consolidation)
    for prof in (results.profiles or []):
        parts.append(f"[Profile] {prof.profile_data}")

    # Past conversation episodes, most relevant first
    for ep in (results.episodes or []):
        parts.append(f"[Past conversation] {ep.episode}")

    # Turns from this session that extraction hasn't folded into an episode yet
    for msg in (results.unprocessed_messages or []):
        parts.append(f"[Just said] {msg.role}: {msg.content}")

    return "\n".join(parts) if parts else "No relevant memories yet."
```

<Note>
  **Two layers of context, and why you want both**

  Episodes are consolidated memory, so they only appear once extraction has run.

  `unprocessed_messages` is the current session's raw buffer, returned when you
  pin one session with `filters={"session_id": ...}`. It covers the window where
  something was just said but hasn't been extracted yet.

  Reading both means a brand-new user still gets continuity from their very first
  turn, instead of the assistant drawing a blank until the first episode lands.

  Searching by `user_id` alone returns episodes only.
</Note>

## Complete assistant loop

```python theme={null}
from everos_cloud import EverOS

client = EverOS()


class PersonalAssistant:
    def __init__(self, user_id: str, session_id: str):
        self.user_id = user_id
        self.session_id = session_id

    def _get_context(self, query: str) -> str:
        results = client.search(
            query, user_id=self.user_id, method="hybrid",
            top_k=5, include_profile=True,
            filters={"session_id": self.session_id},
        )
        parts = [f"[Profile] {p.profile_data}" for p in (results.profiles or [])]
        parts += [f"[Past conversation] {e.episode}" for e in (results.episodes or [])]
        parts += [f"[Just said] {m.role}: {m.content}"
                  for m in (results.unprocessed_messages or [])]
        return "\n".join(parts) if parts else "No relevant memories yet."

    def _generate(self, user_message: str, context: str) -> str:
        prompt = f"""You are a helpful personal assistant. Use the context about
the user to personalize your response. Don't mention memory unless asked.

MEMORY CONTEXT:
{context}

USER MESSAGE:
{user_message}"""
        # Replace with your LLM call (Anthropic, OpenAI, etc.)
        # resp = anthropic.messages.create(model="claude-opus-4-8",
        #     max_tokens=1024, messages=[{"role": "user", "content": prompt}])
        # return resp.content[0].text
        return f"[LLM reply grounded in: {context[:80]}...]"

    def chat(self, user_message: str) -> str:
        # 1. Recall
        context = self._get_context(user_message)
        # 2. Generate
        reply = self._generate(user_message, context)
        # 3. Persist the turn
        client.add(
            session_id=self.session_id,
            messages=[
                {"sender_id": self.user_id, "role": "user", "content": user_message},
                {"sender_id": "assistant", "role": "assistant", "content": reply},
            ],
        )
        return reply


assistant = PersonalAssistant("user_alice", "assistant_session_001")

print(assistant.chat("I prefer meetings in the morning, before 10am."))
print(assistant.chat("What time works best for our call tomorrow?"))
# The second turn recalls the morning-meeting preference from the first.
```

<Tip>
  By default extraction runs on its own schedule. If you want a memory available
  for recall *right now*, for example at the end of a session, call
  `client.flush(session_id)` to force extraction of what's ready.
</Tip>

## Best practices

<AccordionGroup>
  <Accordion title="Keep context focused">
    Limit retrieved memories so you don't overwhelm the LLM context window.

    ```python theme={null}
    top_k=5                 # top-k most relevant episodes
    context = context[:2000]  # optional hard cap (~500 tokens)
    ```
  </Accordion>

  <Accordion title="Choose the right search method">
    ```python theme={null}
    method="vector"   # semantic similarity, paraphrased queries
    method="hybrid"   # keyword + vector + rerank (recommended default)
    method="agentic"  # LLM-guided, for complex multi-part questions (slower)
    ```
  </Accordion>

  <Accordion title="Episodes vs profile">
    Use **episodes** for what happened in past conversations (available within
    seconds). Use the **profile** for stable, consolidated preferences and traits.
    It's built in the background over many interactions, so treat it as
    optional context that gets richer over time.
  </Accordion>
</AccordionGroup>

## Next steps

<CardGroup cols={2}>
  <Card title="Retrieval methods" icon="magnifying-glass" href="/cloud/agentic-retrieval">
    Vector, hybrid, and agentic retrieval in depth.
  </Card>

  <Card title="Python integration" icon="python" href="/cookbook/python-integration">
    Production patterns: error handling, retries, clean client setup.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.