# EverOS > EverOS is the Memory Operating System for Agentic AI. It gives LLM agents persistent, structured memory that extracts knowledge from conversations and multimodal data, resolves contradictions, and retrieves context intelligently — so agents remember, learn, and evolve across sessions. Docs: https://docs.evermind.ai Console & API keys: https://everos.evermind.ai GitHub (open source): https://github.com/EverMind-AI/EverOS Cloud Python SDK: pip install everos-cloud OSS Python package: pip install everos API base URL: https://api.evermind.ai (Cloud) · http://127.0.0.1:8000 (self-hosted default) Auth: Bearer token in Authorization header Contact: contact@evermind.ai | Discord: https://discord.gg/geHdX4F24B ## Unified v2 API The **v2 Memory API** (`/api/v2/memory/*`) is shared across **Cloud** (managed SaaS) and **self-hosted OSS** — same request/response contracts, you just choose where it runs. The v1 Cloud and v1 OSS APIs are legacy, kept for compatibility. What differs between the two: | | Cloud | OSS | |---|---|---| | `add` · `flush` · `get` · `search` | yes | yes | | `edit` · `delete` · memory tags | yes | no | | Async task API (`/api/v2/tasks/*`) | yes | no — writes are direct, so there is nothing to poll | | Knowledge bases | `/api/v2/knowledge_bases/{kb_id}/…` — many named bases, each starting with no categories | `/api/v2/knowledge/…` — one knowledge space per deployment, no `kb_id` in the path | | Multimodal input | upload via `/api/v2/object/sign`, pass the object key as `uri` | pass `uri` (`http(s)://`, `file://`) or inline `base64` | | Reflection (weekly episode consolidation, off by default) | runs on schedule; no trigger endpoint yet | enable in `ome.toml`; `POST /api/v2/ome/trigger` runs it by hand | ## TL;DR (2 API Calls) ``` 1. POST /api/v2/memory/add — Add messages (session_id, messages[] each with sender_id) 2. POST /api/v2/memory/search — Retrieve context (query, user_id or agent_id, method="hybrid", top_k=5) ``` After these 2 calls, your agent has persistent, cross-session memory. EverOS replaces ad-hoc memory (chat history, RAG hacks) with a structured, persistent memory layer. Note: `add` is asynchronous by default — it returns `status: "queued"` (HTTP 202) and extraction runs in the background on its own. You do NOT need to trigger it. Pass `async_mode: false` for a synchronous write that returns the engine's own result (`"accumulated"` or `"extracted"`). `POST /api/v2/memory/flush` forces extraction for a session that is still open, but it only sees messages that have already landed — called right after a queued `add` it returns `"no_extraction"`, which means "nothing pending", not an error. Reads: `get` sees a memory as soon as it is extracted; `search` reads a vector index that lags extraction by seconds. ## Recommended Agent Loop ``` On every user interaction: 1. SEARCH → POST /api/v2/memory/search (user_id=..., method="hybrid", top_k=5-10) Use returned episode summaries and profile attributes as context (not raw message logs). If no memories are found, proceed without memory context. 2. GENERATE → Call your LLM with the search results (episode summaries + profile attributes) as context. 3. ADD → POST /api/v2/memory/add (store the user + assistant messages, each with sender_id) Extraction runs in the background on its own — no follow-up call is required. Only reach for /flush when you need a still-open session extracted right now, and only after the add has landed. Tip: Use consistent session_id values to group related interactions and avoid fragmented memory. Repeat across sessions. EverOS handles consolidation and profile evolution automatically. ``` ## When to Use EverOS Reach for it when memory must survive the session — user modelling, preferences and history, automatic extraction from conversation, per-participant attribution in multi-party chats (`sender_id`). If you only need static document retrieval with no persistence between sessions, plain RAG is the simpler tool. ## How It Works Three phases: **Formation** (conversation is segmented and distilled into discrete memories), **Consolidation** (those memories keep merging, correcting and re-profiling themselves in the background), **Recollection** (a query assembles context from them rather than matching text). The middle one is what makes recall improve over time. Mechanically: Messages you `add` accumulate in a session buffer. A detector watches for a topic shift or a time gap; when one trips, the buffer closes into a **MemCell** and extraction produces an episode narrative plus the atomic facts inside it. That is why `add` returns before anything exists, why `flush` on a session with nothing pending answers `no_extraction`, and why a memory appears seconds after the turn that produced it. Two places the MemCell is visible to you: an episode's `parent_id` points at the one it came from, and **Cloud quota is counted in MemCells** — roughly one per ten raw messages. Afterwards, consolidation runs on its own: duplicates merge, contradictions between old and new facts resolve, and each user's profile is rewritten as it learns more. Nothing you call triggers it, which is why the same query can answer better a day later — and why `edit` expects a profile extraction has already built. ## Memory Types `memory_type` (in get/search) is one of: | Type | `memory_type` | What it captures | |------|---------------|------------------| | Episode | `episode` | Narrative summaries of conversations, each carrying the atomic facts extracted from it | | Profile | `profile` | Persistent user attributes and preferences | | Agent Case | `agent_case` | Task approach and quality from agent trajectories | | Agent Skill | `agent_skill` | Generalized skills distilled from agent cases | Agent cases + skills together are what an agent has learned about itself. ## Cloud SDK Quickstart ```python from everos_cloud import EverOS client = EverOS(api_key="your_api_key") # required: 1.x reads no env vars; pass host= for a non-prod gateway # Add messages — sender_id is who the memory belongs to. The SDK stamps timestamps # with the current time (ms) when omitted; over raw HTTP `timestamp` is required. client.add( session_id="session_001", messages=[ {"sender_id": "user_001", "role": "user", "content": "I like black coffee, no sugar."}, {"sender_id": "assistant", "role": "assistant", "content": "Got it, I'll remember that."}, ], ) # Extraction runs on its own; no flush needed. For a deterministic write: # client.add(..., async_mode=False) # Search — requires user_id or agent_id results = client.search("coffee preference", user_id="user_001", method="hybrid", top_k=5) # results.episodes[] — ranked episode hits: # .episode full narrative → paste into LLM prompt as context # .summary short summary (~200 chars) # .atomic_facts[] single-sentence facts (.content) # results.profiles[] — populated when include_profile=True # Extraction is async, so the newest turns may not be in an episode yet. Pin a session # to also get that in-flight tail back as raw messages: results = client.search("coffee preference", user_id="user_001", filters={"session_id": "session_001"}) # results.unprocessed_messages[] — raw buffered messages, oldest first, no score. # Only loads for a top-level session_id equality scalar (not inside AND/OR, not {"eq": ...}). # Get the profile directly profile = client.get("profile", user_id="user_001") # profile.profiles[0].profile_data — dict of user attributes and preferences ``` Each method returns the response `.data`. The typed low-level client is available as `client.memory` / `client.storage`. ## v2 API Endpoints | Operation | Method & Path | Availability | Key params | |-----------|---------------|--------------|------------| | Add messages | `POST /api/v2/memory/add` | OSS + Cloud | `session_id`, `messages[]` (each `sender_id`, `role`, `timestamp`, `content` — all required over HTTP) | | Force extraction | `POST /api/v2/memory/flush` | OSS + Cloud | `session_id` | | Get memories | `POST /api/v2/memory/get` | OSS + Cloud | `memory_type`, `user_id` or `agent_id`, `page`, `page_size` | | Search memories | `POST /api/v2/memory/search` | OSS + Cloud | `query`, `user_id` or `agent_id`, `method`, `top_k` · add `filters.session_id` to also get `unprocessed_messages[]` (the not-yet-extracted tail) | | Edit profile | `POST /api/v2/memory/edit` | Cloud-only | `user_id`, `operations[]` (add/update/delete profile items) · requires a profile extraction already built | | Delete memories | `POST /api/v2/memory/delete` | Cloud-only | scope by `user_id` / `agent_id` / `session_id` | | Upload multimodal (pre-sign) | `POST /api/v2/object/sign` | Cloud (Storage) | `objectList[]` | | Tag memories | `POST /api/v2/memory/tag/{bind,unbind,replace}` | Cloud | `memory_type`, `memory_ids[]`, `tags[]` · bind adds, unbind removes, replace overwrites (empty list clears) | | Create / list knowledge base | `POST` / `GET /api/v2/knowledge_bases` | Cloud | `name`, `description` | | Read / update / delete knowledge base | `GET` / `PATCH` / `DELETE /api/v2/knowledge_bases/{kb_id}` | Cloud | delete cascades to documents, topics and index | | Ingest a document | `POST /api/v2/knowledge_bases/{kb_id}/documents` | Cloud | `title`, `content` (inline text or an object key as `uri`), optional `category_id` · async: 202 + `task_id`, no document id | | Browse documents | `GET .../documents` · `GET .../documents/{doc_id}` | Cloud | `topic_count` > 0 means ingest finished | | Search a knowledge base | `POST /api/v2/knowledge_bases/{kb_id}/search` | Cloud | `query`, `method`, `top_k`, `include: ["content"]` · hits are topics, scored per-response | | Read topics | `GET .../documents/{doc_id}/topics[/{topic_id}]` | Cloud | flat DFS order; includes one synthetic root, so `len(topics) == topic_count + 1` | | Categories | `POST` / `GET .../categories` · `PATCH` / `DELETE .../categories/{id}` | Cloud | deleting one reassigns its documents to uncategorized | | Async task status | `GET /api/v2/tasks/{task_id}` · `GET /api/v2/tasks` · `GET /api/v2/tasks/stats` | Cloud | the completion side of every async write | `app_id` / `project_id` both default to `"default"` and partition memory; queries never cross scopes. ## Rules That Will 4xx You Every one of these is enforced server-side and returns the message shown. Get them right and the API is quiet. | Rule | What you get if you break it | |---|---| | `timestamp` on every message, **unix milliseconds** | omitted → `400` `` `timestamp` is required ``; seconds-scale → `422` `` `timestamp` must be a unix millisecond timestamp `` | | `get` / `search`: **exactly one** of `user_id` / `agent_id` | neither or both → `422` `exactly one of user_id / agent_id is required` | | `get`: `memory_type` must match the owner — a user owns `episode` and `profile`, an agent owns `agent_case` and `agent_skill` | `422` `memory_type 'agent_case' is not valid when user_id is set` | | `delete`: at least one of `user_id` / `agent_id` / `session_id` | empty body → `422` `at least one of user_id / agent_id / session_id required` | | `top_k`: `-1` or 1–100 | `422` `top_k must be -1 or in 1..100` | Other caps, all rejected with `422` when exceeded: 1–500 messages per `add`, `session_id` 1–128 chars, 1–50 operations per `edit`, page size 1–100, and for tags up to 200 memory ids × 100 tags with each tag ≤32 chars. Every successful response is `{"request_id": "...", "data": {...}}`. Read your result out of `data`. ## Knowledge Bases (Cloud) A document library that lives beside memory, retrieved at the **topic** level — an LLM splits each document into topics, and search returns those. The one thing to get right: **ingest is asynchronous and its ack carries no document id.** The id is minted downstream, so you poll the task, then find the document by title. ```python kb = client.kb_create("Employee Handbook") ack = client.doc_ingest(kb.id, "Leave policy", "Employees accrue 20 days...") task = client.task_wait(ack.task_id, timeout=300) # ack.task_id, not a doc id doc = next(d for d in client.doc_list(kb.id).documents if d.title == "Leave policy") # doc.topic_count > 0 is the authoritative "ingest finished" signal hits = client.kb_search(kb.id, "how much leave do I get", top_k=5) for h in hits.hits: print(h.name, h.summary, h.document.title) ``` Four things that surprise callers: - A new knowledge base has **no categories**. Until you create some, every document stays uncategorized and the classifier has nothing to choose from. - Results include the synthetic **document-root topic** (`depth` 0, the document's own title). It has no body — skip `depth == 0` if you only want real sections. - `hit.score` is normalized **within each response**, so it compares inside one result set and nowhere else. The top hit sits near the top of the range however weak the pool is, which makes a fixed `score_threshold` a relative cut rather than a relevance bar. - Topic bodies are omitted unless you pass `include: ["content"]`. ## Retrieval Methods | Method | Latency | Best for | |--------|---------|----------| | `keyword` | <100ms | Exact terms, known phrases (BM25) | | `vector` | 200-500ms | Semantic similarity, paraphrased queries | | `hybrid` | 200-600ms | **Recommended default** (keyword + vector + rerank) | | `agentic` | 2-5s | Complex multi-part questions (LLM-guided). Use only when hybrid is insufficient. Fallback to hybrid on timeout. | Start with `method="hybrid"`, `top_k=5` for chat or `top_k=10` for research/analysis. `top_k` is either `-1` (the default — the server decides how many to return) or 1–100; anything else is rejected. Only escalate to `agentic` for complex, multi-part queries. ## Filters DSL Optional on `get` and `search`, alongside the top-level `user_id` / `agent_id`: ```json {"AND": [{"timestamp": {"gte": 1700000000000}}, {"session_id": "s1"}]} ``` Filterable fields — anything else is rejected with `422`: | field | matches | |---|---| | `session_id` | the session the memory came from | | `sender_id` | one participant — matched against the episode's `sender_ids` | | `timestamp` | when the memory happened | | `parent_id` / `parent_type` | what a memory derives from: an episode's parent is its memcell, an atomic fact's parent is its episode | | `tag` | a tag you attached — **Cloud only, and `get` only**; passing it to `search` is a `422`, and the OSS allow-list has no `tag` field at all | Operators: `eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `in`. A bare value is `eq` shorthand. `AND` / `OR` nest freely. `user_id`, `agent_id`, `app_id` and `project_id` do **not** belong inside `filters` — they are top-level request fields, and putting them in a filter is a `422`. ## Multimodal Support Messages accept an array of content items instead of a plain string. Supported types: `text`, `image`, `audio`, `doc`, `pdf`, `html`, `email`. Non-text files are preprocessed to text (image OCR, audio transcription, PDF/Office parsing) before extraction. - **Cloud SDK**: `object_key = client.upload("photo.jpg")` presigns and uploads in one call, returning the key; reference it in a message as `{"type": "image", "uri": object_key}`. - **Raw API**: `POST /api/v2/object/sign` → POST the bytes to the returned URL with the fields it gives you → use the returned `objectKey` as `uri`. - A non-text item that carries only `base64` and no `uri` is **skipped** by the parse step on Cloud — upload it and pass the key. - **OSS**: pass `uri` directly (`http://`, `https://`, or `file://` for server-local paths) or `base64` inline. Requires a vision/audio LLM configured. ## Further Reading For full API schemas, request/response examples, and open-source deployment: - llms-full.txt: detailed reference (same directory) - Full docs: https://docs.evermind.ai - API Reference: https://docs.evermind.ai/api-reference/introduction - Cloud Quickstart: https://docs.evermind.ai/cloud/quickstart