---
title: memory-atlas
canonical_url: "https://www.modelscope.cn/studios/vqhvqh/memory-atlas"
md_url: "https://www.modelscope.cn/studios/vqhvqh/memory-atlas.md"
repository: vqhvqh/memory-atlas
last_updated: 2026-03-26
sdk_type: gradio
sdk_version: 6.2.0
downloads: 0
stars: 0
---

# memory-atlas

> memory-atlas - vqhvqh 在 ModelScope 创建的在线 Demo。MemoryAtlas 是一个面向 AI 应用开发者的 Python SDK，为 LangChain 1.0 agent 提供透明的记忆管理能力。现有记忆系统（Mem0、Zep、Letta 等）都是"被动检索"模式——query 来了，搜索，返回结果。MemoryAtlas 的核心差异是引入游戏引擎的"主动场景管理"——像游戏引擎管理纹理和模型一样，主动预测、预加载、分级、剔除记忆，让 agent…

vqhvqh/memory-atlas 是 ModelScope 魔搭社区上的在线可交互 Demo（创空间），基于 gradio 6.2.0 构建。

- **Repository**: vqhvqh/memory-atlas
- **SDK**: gradio
- **SDK version**: 6.2.0
- **Downloads**: 0
- **Stars**: 0
- **Last updated**: 2026-03-26

Source: https://www.modelscope.cn/studios/vqhvqh/memory-atlas

---

# MemoryAtlas

> The first AI agent memory system designed with game engine resource management principles.

<p align="center">
  <a href="#installation">Installation</a> •
  <a href="#quick-start">Quick Start</a> •
  <a href="#core-idea">Core Idea</a> •
  <a href="#cli">CLI</a> •
  <a href="#benchmarks">Benchmarks</a> •
  <a href="docs/API.md">API Docs</a> •
  <a href="docs/README_CN.md">中文文档</a>
</p>

---

## Why MemoryAtlas?

Existing memory systems (Mem0, Zep, Letta) are all **passive retrieval** — query comes in, search, return results.

MemoryAtlas introduces **active scene management** from game engines — proactively predicting, preloading, leveling, and culling memories so the agent always has the optimal memory view.

| Feature | Mem0 | Letta | Zep | MemoryAtlas |
|---|:---:|:---:|:---:|:---:|
| Predictive prefetching | ✗ | ✗ | ✗ | ✓ |
| Frustum culling (active unload) | ✗ | ✗ | ✗ | ✓ |
| Multi-level LOD | ✗ | Partial | ✗ | ✓ |
| Three-tier cache | ✗ | ✗ | ✗ | ✓ |
| Memory clusters (Asset Bundling) | ✗ | ✗ | ✗ | ✓ |
| Zero external dependencies | ✗ | ✗ | ✗ | ✓ |
| Local storage, Git-friendly | ✗ | ✗ | ✗ | ✓ |

## Installation

```bash
pip install memory-atlas

# Local embeddings (works offline):
pip install memory-atlas[local-embedding]
```

## Quick Start

### One-line LangChain Integration

```python
from memory_atlas.langchain import MemoryAtlasMiddleware

memory = MemoryAtlasMiddleware(
    storage_path="./my_agent_memory",
    embedding_model="local",
    max_memory_tokens=2000,
)

agent = create_agent(
    model="openai:gpt-4o",
    tools=[...],
    middleware=[memory],  # That's it.
)
```

### Standalone Usage

```python
from memory_atlas import MemoryEngine

engine = MemoryEngine(storage_path="./memory", embedding_model="local")

# Ingest
engine.ingest("JWT refresh token has a race condition. Decided to use sliding window strategy.")

# Retrieve (scene manager handles caching, LOD, prefetching automatically)
memories = engine.retrieve("token expiration issue")
for m in memories:
    print(f"[{m.lod}] {m.display_text}")

# Expand to full content
detail = engine.expand(memories[0].id)

# Export / Import
engine.export_memories("backup.json")
engine.import_memories("backup.json", mode="merge")

engine.close()
```

### LlamaIndex / CrewAI

```python
# LlamaIndex
from memory_atlas.integrations.llamaindex import MemoryAtlasRetriever
retriever = MemoryAtlasRetriever(storage_path="./memory")
nodes = retriever.retrieve("auth token bug")

# CrewAI
from memory_atlas.integrations.crewai import MemoryAtlasTool
tool = MemoryAtlasTool(storage_path="./memory")
result = tool.run("search for auth bugs")
```

## CLI

```bash
memory-atlas init -s ./my_memory          # Initialize memory store
memory-atlas ingest "JWT race condition" -s ./my_memory  # Ingest content
memory-atlas search "token expiry" -s ./my_memory        # Search memories
memory-atlas stats -s ./my_memory                        # Show statistics
memory-atlas forget -s ./my_memory                       # Forget low-activity memories
memory-atlas export backup.json -s ./my_memory           # Export to JSON
memory-atlas import backup.json -s ./my_memory           # Import from JSON
memory-atlas clusters -s ./my_memory                     # List memory clusters
```

## Core Idea

Game engines face the same problem as AI agents: **the world is huge, but the viewport is limited.**

Game engines solved this decades ago with LOD, prefetching, frustum culling, and layered caching. MemoryAtlas maps these patterns directly to memory management.

```
┌─────────────────────────────────────────────────────────────┐
│                      Scene Manager                          │
│                                                             │
│  ┌───────────────┐  ┌───────────────┐  ┌────────────────┐  │
│  │  Prefetcher   │  │  LOD Manager  │  │ Frustum Culler │  │
│  │               │  │               │  │                │  │
│  │  Predict next │  │  Far → L0     │  │  Topic shift → │  │
│  │  turn's needs │  │  Near → L1    │  │  actively      │  │
│  │  & preload    │  │  Focus → L2   │  │  unload        │  │
│  └───────┬───────┘  └───────┬───────┘  └───────┬────────┘  │
│          └──────────────────┼──────────────────┘           │
│                             ▼                               │
│              Memory View (optimal context at all times)     │
└─────────────────────────────────────────────────────────────┘
```

| Game Engine Concept | MemoryAtlas Equivalent |
|---|---|
| LOD (Level of Detail) | L0 label / L1 summary / L2 full content |
| Streaming Prefetch | Predict next topics, preload to warm cache |
| Frustum Culling | Detect topic shifts, demote irrelevant memories |
| Asset Bundling | Memory clusters, load/unload as a unit |
| Layered Cache | Hot (in-memory) → Warm (LRU) → Cold (DuckDB + files) |

### Three-tier Cache

```
  Hot   ← Active context, O(1) access, ~0.3µs
  Warm  ← Preloaded candidates + recently demoted, LRU
  Cold  ← DuckDB vector search + tree reasoning, ~22ms
```

### Memory Clusters (Auto-bundling)

When an entity accumulates enough linked memories, they're automatically bundled:

```python
engine.ingest("JWT token expiration bug")       # jwt → 1 memory
engine.ingest("JWT refresh token race condition") # jwt → 2 memories
engine.ingest("JWT sliding window strategy")     # jwt → 3 memories → auto-creates cluster:jwt
```

### Forgetting

```
activity = importance × e^(-λ × days) × log(access_count + 2)
```

Low-activity memories are automatically compressed (drop L2, keep L1) or archived (keep only L0 label).

## Architecture

```
┌──────────────────────────────────────────────────────────────┐
│                      MemoryAtlas SDK                         │
│                                                              │
│  ┌────────────────────────────────────────────────────────┐  │
│  │           Scene Manager (core differentiator)          │  │
│  │  Prefetcher ── LOD Manager ── Frustum Culler           │  │
│  └────────────────────────┬───────────────────────────────┘  │
│                           │                                  │
│  ┌────────────┐  ┌────────┴───────┐  ┌─────────────────┐    │
│  │ Ingestion  │  │   Retrieval    │  │  Cache Manager  │    │
│  │ Pipeline   │  │   Engine       │  │  Hot/Warm/Cold  │    │
│  │ chunk/     │  │ vector+tree    │  │                 │    │
│  │ extract/   │  │ → fusion       │  │  Cluster Mgr    │    │
│  │ summarize  │  │                │  │  (auto-bundle)  │    │
│  └────────────┘  └────────────────┘  └─────────────────┘    │
│                                                              │
│  ┌───────────────────────────────────────────────────────┐   │
│  │  Storage: DuckDB (index) + Markdown (L2 content)      │   │
│  └───────────────────────────────────────────────────────┘   │
│                                                              │
│  ┌───────────────────────────────────────────────────────┐   │
│  │  Integrations: LangChain │ LlamaIndex │ CrewAI │ CLI  │   │
│  └───────────────────────────────────────────────────────┘   │
└──────────────────────────────────────────────────────────────┘
```

## Benchmarks

| Metric | Result | Target |
|---|---|---|
| Cache hit rate | **76%** | > 60% ✅ |
| Prefetch accuracy | **100%** | > 50% ✅ |
| Token savings (LOD) | **93.4%** | > 40% ✅ |
| Cache retrieval latency | **0.3µs** | < 10ms ✅ |
| Cold retrieval latency | **22ms** | < 200ms ✅ |
| Cache speedup | **69,000x** | — |

```bash
uv run python -m benchmarks.cache_hit_rate
uv run python -m benchmarks.prefetch_accuracy
uv run python -m benchmarks.token_savings
uv run python -m benchmarks.latency_comparison
```

## Tech Stack

| Component | Choice | Why |
|---|---|---|
| Metadata index | DuckDB | Embedded, zero deps, vector ops |
| Memory storage | Markdown | Human-readable, Git-friendly |
| LLM interface | LiteLLM | Vendor-agnostic |
| Embedding | sentence-transformers / OpenAI / Ollama | Local + cloud |
| CLI | Typer | Type-safe, auto-generated help |
| Package manager | uv | Fast |

## Project Structure

```
src/memory_atlas/
├── engine.py              # Core engine
├── config.py              # Configuration
├── cli.py                 # CLI tool
├── scene/                 # ⭐ Scene Manager
│   ├── manager.py         #   Orchestrates prefetch/cull/LOD
│   ├── prefetcher.py      #   Predictive preloading
│   ├── culler.py          #   Frustum culling
│   └── lod.py             #   Precision management
├── core/
│   ├── registry.py        #   DuckDB metadata
│   ├── tree_index.py      #   Semantic tree index
│   └── cluster.py         #   Memory clusters
├── ingestion/             # Ingestion pipeline
├── retrieval/             # Retrieval engine (vector + tree + fusion)
├── storage/               # File store + three-tier cache
├── maintenance/           # Forgetting mechanism
├── llm/                   # LLM + Embedding providers
├── langchain/             # LangChain integration
└── integrations/          # LlamaIndex / CrewAI
```

## Development

```bash
# Install dependencies
uv sync --dev

# Run tests (91)
uv run python -m pytest tests/ -v

# Lint
uv run ruff check src/ tests/

# Run benchmarks
uv run python -m benchmarks.cache_hit_rate
```

## License

MIT
